Searcharxiv⌕ Search

arXiv subjects

Sayan Ghosh

Publications and source records attributed to Sayan Ghosh.

At least 73 records · Page 4Linked to original sources

Reconciling Security and Communication Efficiency in Federated Learning

Cross-device Federated Learning is an increasingly popular machine learning setting to train a model by leveraging a large population of client devices with high privacy and security guarantees. However, communication efficiency remains a major bottleneck when scaling federated learning to production environments, particularly due to bandwidth constraints during uplink communication. In this paper, we formalize and address the problem of compressing client-to-server model updates under the Secure Aggregation primitive, a core component of Federated Learning pipelines that allows the server to aggregate the client updates without accessing them individually. In particular, we adapt standard scalar quantization and pruning methods to Secure Aggregation and propose Secure Indexing, a variant of Secure Aggregation that supports quantization for extreme compression. We establish state-of-the-art results on LEAF benchmarks in a secure Federated Learning setup with up to 40$\times$ compression in uplink communication with no meaningful loss in utility compared to uncompressed baselines.

cs.LG↗

ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding

While large language models have shown exciting progress on several NLP benchmarks, evaluating their ability for complex analogical reasoning remains under-explored. Here, we introduce a high-quality crowdsourced dataset of narratives for employing proverbs in context as a benchmark for abstract language understanding. The dataset provides fine-grained annotation of aligned spans between proverbs and narratives, and contains minimal lexical overlaps between narratives and proverbs, ensuring that models need to go beyond surface-level reasoning to succeed. We explore three tasks: (1) proverb recommendation and alignment prediction, (2) narrative generation for a given proverb and topic, and (3) identifying narratives with similar motifs. Our experiments show that neural language models struggle on these tasks compared to humans, and these tasks pose multiple learning challenges.

cs.CL↗

CLUES: A Benchmark for Learning Classifiers using Natural Language Explanations

Supervised learning has traditionally focused on inductive learning by observing labeled examples of a task. In contrast, humans have the ability to learn new concepts from language. Here, we explore training zero-shot classifiers for structured data purely from language. For this, we introduce CLUES, a benchmark for Classifier Learning Using natural language ExplanationS, consisting of a range of classification tasks over structured data along with natural language supervision in the form of explanations. CLUES consists of 36 real-world and 144 synthetic classification tasks. It contains crowdsourced explanations describing real-world tasks from multiple teachers and programmatically generated explanations for the synthetic tasks. To model the influence of explanations in classifying an example, we develop ExEnt, an entailment-based model that learns classifiers using explanations. ExEnt generalizes up to 18% better (relative) on novel tasks than a baseline that does not use explanations. We delineate key challenges for automated learning from explanations, addressing which can lead to progress on CLUES in the future. Code and datasets are available at: https://clues-benchmark.github.io.

cs.CL↗

Learning to Mediate Disparities Towards Pragmatic Communication

Human communication is a collaborative process. Speakers, on top of conveying their own intent, adjust the content and language expressions by taking the listeners into account, including their knowledge background, personalities, and physical capabilities. Towards building AI agents with similar abilities in language communication, we propose Pragmatic Rational Speaker (PRS), a framework extending Rational Speech Act (RSA). The PRS attempts to learn the speaker-listener disparity and adjust the speech accordingly, by adding a light-weighted disparity adjustment layer into working memory on top of speaker's long-term memory system. By fixing the long-term memory, the PRS only needs to update its working memory to learn and adapt to different types of listeners. To validate our framework, we create a dataset that simulates different types of speaker-listener disparities in the context of referential games. Our empirical results demonstrate that the PRS is able to shift its output towards the language that listener are able to understand, significantly improve the collaborative task outcome.

cs.CL↗

Simulation of Nuclear Recoils due to Supernova Neutrino-induced Neutrons in Liquid Xenon Detectors

Neutrinos from supernova (SN) bursts can give rise to detectable number of nuclear recoil (NR) events through the coherent elastic neutrino-nucleus scattering (CE$ν$NS) process in large scale liquid xenon detectors designed for direct dark matter search, depending on the SN progenitor mass and distance. Here we show that in addition to the direct NR events due to CE$ν$NS process, the SN neutrinos can give rise to additional nuclear recoils due to the elastic scattering of neutrons produced through inelastic interaction of the neutrinos with the xenon nuclei. We find that the contribution of the supernova neutrino-induced neutrons ($ν$I$n$) can significantly modify the total xenon NR spectrum at large recoil energies compared to that expected from the CE$ν$NS process alone. Moreover, for recoil energies $\gtrsim20$ keV, dominant contribution is obtained from the ($ν$I$n$) events. We numerically calculate the observable S1 and S2 signals due to both CE$ν$NS and $ν$I$n$ processes for a typical liquid xenon based detector, accounting for the multiple scattering effects of the neutrons in the case of $ν$I$n$, and find that sufficiently large signal events, those with S1$\gtrsim$50 photo-electrons (PE) and S2$\gtrsim$2300 PE, come mainly from the $ν$I$n$ scatterings.

astro-ph.HE↗

Reinforcement Learning based Sequential Batch-sampling for Bayesian Optimal Experimental Design

Engineering problems that are modeled using sophisticated mathematical methods or are characterized by expensive-to-conduct tests or experiments, are encumbered with limited budget or finite computational resources. Moreover, practical scenarios in the industry, impose restrictions, based on logistics and preference, on the manner in which the experiments can be conducted. For example, material supply may enable only a handful of experiments in a single-shot or in the case of computational models one may face significant wait-time based on shared computational resources. In such scenarios, one usually resorts to performing experiments in a manner that allows for maximizing one's state-of-knowledge while satisfying the above mentioned practical constraints. Sequential design of experiments (SDOE) is a popular suite of methods, that has yielded promising results in recent years across different engineering and practical problems. A common strategy, that leverages Bayesian formalism is the Bayesian SDOE, which usually works best in the one-step-ahead or myopic scenario of selecting a single experiment at each step of a sequence of experiments. In this work, we aim to extend the SDOE strategy, to query the experiment or computer code at a batch of inputs. To this end, we leverage deep reinforcement learning (RL) based policy gradient methods, to propose batches of queries that are selected taking into account entire budget in hand. The algorithm retains the sequential nature, inherent in the SDOE, while incorporating elements of reward based on task from the domain of deep RL. A unique capability of the proposed methodology is its ability to be applied to multiple tasks, for example optimization of a function, once its trained. We demonstrate the performance of the proposed algorithm on a synthetic problem, and a challenging high-dimensional engineering problem.

cs.LG↗

Leggett-Garg inequality in Markovian quantum dynamics: role of temporal sequencing of coupling to bath

We study Leggett-Garg inequalities (LGIs) for a two level system (TLS) undergoing Markovian dynamics described by unital maps. We find analytic expression of LG parameter $K_{3}$ (simplest variant of LGIs) in terms of the parameters of two distinct unital maps representing time evolution for intervals: $t_{1}$ to $t_{2}$ and $t_{2}$ to $t_{3}$. We show that the maximum violation of LGI for these maps can never exceed well known Lüders bound of $K_{3}^{L\ddot{u}ders}=3/2$ over the full parameter space. We further show that if the map for the time interval $t_{1}$ to $t_{2}$ is non-unitary unital then irrespective of the choice of the map for interval $t_{2}$ to $t_{3}$ we can never reach Lüders bound. On the other hand, if the measurement operator eigenstates remain pure upon evolution from $t_{1}$ to $t_{2}$, then depending on the degree of decoherence induced by the unital map for the interval $t_{2}$ to $t_{3}$ we may or may not obtain Lüders bound. Specifically, we find that if the unital map for interval $t_{2}$ to $t_{3}$ leads to the shrinking of the Bloch vector beyond half of its unit length, then achieving the bound $K_{3}^{L\ddot{u}ders}$ is not possible. Hence our findings not only establish a threshold for decoherence which will allow for $K_{3} = K_{3}^{L\ddot{u}ders}$, but also demonstrate the importance of temporal sequencing of the exposure of a TLS to Markovian baths in obtaining Lüders bound.

quant-ph↗

Mapping Language to Programs using Multiple Reward Components with Inverse Reinforcement Learning

Mapping natural language instructions to programs that computers can process is a fundamental challenge. Existing approaches focus on likelihood-based training or using reinforcement learning to fine-tune models based on a single reward. In this paper, we pose program generation from language as Inverse Reinforcement Learning. We introduce several interpretable reward components and jointly learn (1) a reward function that linearly combines them, and (2) a policy for program generation. Fine-tuning with our approach achieves significantly better performance than competitive methods using Reinforcement Learning (RL). On the VirtualHome framework, we get improvements of up to 9.0% on the Longest Common Subsequence metric and 14.7% on recall-based metrics over previous work on this framework (Puig et al., 2018). The approach is data-efficient, showing larger gains in performance in the low-data regime. Generated programs are also preferred by human evaluators over an RL-based approach, and rated higher on relevance, completeness, and human-likeness.

cs.CL↗

Detecting Cross-Geographic Biases in Toxicity Modeling on Social Media

Online social media platforms increasingly rely on Natural Language Processing (NLP) techniques to detect abusive content at scale in order to mitigate the harms it causes to their users. However, these techniques suffer from various sampling and association biases present in training data, often resulting in sub-par performance on content relevant to marginalized groups, potentially furthering disproportionate harms towards them. Studies on such biases so far have focused on only a handful of axes of disparities and subgroups that have annotations/lexicons available. Consequently, biases concerning non-Western contexts are largely ignored in the literature. In this paper, we introduce a weakly supervised method to robustly detect lexical biases in broader geocultural contexts. Through a case study on a publicly available toxicity detection model, we demonstrate that our method identifies salient groups of cross-geographic errors, and, in a follow up, demonstrate that these groupings reflect human judgments of offensive and inoffensive language in those geographic contexts. We also conduct analysis of a model trained on a dataset with ground truth labels to better understand these biases, and present preliminary mitigation experiments.

cs.CL↗

Adversarial Scrubbing of Demographic Information for Text Classification

Contextual representations learned by language models can often encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated target task. We aim to scrub such undesirable attributes and learn fair representations while maintaining performance on the target task. In this paper, we present an adversarial learning framework "Adversarial Scrubber" (ADS), to debias contextual representations. We perform theoretical analysis to show that our framework converges without leaking demographic information under certain conditions. We extend previous evaluation techniques by evaluating debiasing performance using Minimum Description Length (MDL) probing. Experimental evaluations on 8 datasets show that ADS generates representations with minimal information about demographic attributes while being maximally informative about the target task.

cs.CL↗

Inverse Aerodynamic Design of Gas Turbine Blades using Probabilistic Machine Learning

One of the critical components in Industrial Gas Turbines (IGT) is the turbine blade. Design of turbine blades needs to consider multiple aspects like aerodynamic efficiency, durability, safety and manufacturing, which make the design process sequential and iterative.The sequential nature of these iterations forces a long design cycle time, ranging from several months to years. Due to the reactionary nature of these iterations, little effort has been made to accumulate data in a manner that allows for deep exploration and understanding of the total design space. This is exemplified in the process of designing the individual components of the IGT resulting in a potential unrealized efficiency. To overcome the aforementioned challenges, we demonstrate a probabilistic inverse design machine learning framework (PMI), to carry out an explicit inverse design. PMI calculates the design explicitly without excessive costly iteration and overcomes the challenges associated with ill-posed inverse problems. In this work, the framework will be demonstrated on inverse aerodynamic design of three-dimensional turbine blades.

eess.SP↗

A Fully Bayesian Gradient-Free Supervised Dimension Reduction Method using Gaussian Processes

Modern day engineering problems are ubiquitously characterized by sophisticated computer codes that map parameters or inputs to an underlying physical process. In other situations, experimental setups are used to model the physical process in a laboratory, ensuring high precision while being costly in materials and logistics. In both scenarios, only limited amount of data can be generated by querying the expensive information source at a finite number of inputs or designs. This problem is compounded further in the presence of a high-dimensional input space. State-of-the-art parameter space dimension reduction methods, such as active subspace, aim to identify a subspace of the original input space that is sufficient to explain the output response. These methods are restricted by their reliance on gradient evaluations or copious data, making them inadequate to expensive problems without direct access to gradients. The proposed methodology is gradient-free and fully Bayesian, as it quantifies uncertainty in both the low-dimensional subspace and the surrogate model parameters. This enables a full quantification of epistemic uncertainty and robustness to limited data availability. It is validated on multiple datasets from engineering and science and compared to two other state-of-the-art methods based on four aspects: a) recovery of the active subspace, b) deterministic prediction accuracy, c) probabilistic prediction accuracy, and d) training time. The comparison shows that the proposed method improves the active subspace recovery and predictive accuracy, in both the deterministic and probabilistic sense, when only few model observations are available for training, at the cost of increased training time.

stat.ML↗

Measurements of gamma ray, cosmic muon and residual neutron background fluxes for rare event search experiments at an underground laboratory

Ambient radiation background contributed by the penetrating cosmic ray particles and the radionuclides present in the rock materials have been measured at an underground laboratory located inside a mine at 555 m depth. The laboratory is being set up to explore rare event search processes, such as direct dark matter search, neutrinoless double beta decay, axion search, supernova neutrino detection, etc., that require specific knowledge of the nature and extent of the radiation environment in order to assess the sensitivity reach and also to plan for its reduction for the targeted experiment. The gamma ray background, which is mostly contributed by the primordial radionuclides and their decay chain products, have been measured inside the laboratory and found to be dominated by rock radioactivity for $E_γ\lesssim 3 \,{\rm MeV}$. Shielding of these residual gamma rays for the experiment was also evaluated. The cosmic muon flux, measured inside the laboratory using large area plastic scintillator telescope, was found to be: $(2.051 \pm 0.142 \pm 0.009) \times 10^{-7}\, {\rm cm}^{-2}.{\rm sec}^{-1}$, which agrees reasonably well with simulation results. The neutron background flux has been measured for the radiogenic neutrons and found to be: $(1.61 \pm 0.03) \times 10^{-4} \, {\rm cm}^{-2}.{\rm sec}^{-1}$ for no threshold cut. Detailed GEANT4 simulation for the radiogenic neutrons and the cosmogenic neutrons have been carried out. Effects of multiple scattering of both the types of neutrons within the surrounding rock and the cavern walls were studied and the results for the radiogenic neutrons are found to be in reasonable agreement with experimental results. Neutron fluxes contributed by those neutrons of cosmogenic origin have been reported as function of the energy threshold.

astro-ph.IM↗

Leptophilic-portal Dark Matter in the Light of AMS-02 positron excess

We revisit dark matter annihilation as an explanation of the positron excess reported recently by the AMS-02 satellite-borne experiment. To this end, we propose a particle dark matter model by considering a Two Higgs Doublet Model (2HDM) extended with an additional singlet boson and a singlet fermion. The additional (light) boson mixes with the pseudoscalar inherent in the 2HDM, and the singlet fermion, which is the dark matter candidate, annihilates via this bosonic portal. The dark matter candidate is made leptophilic by choosing the lepton-specific 2HDM and a suitable high value of $\tanβ$. We identify the model parameter space which explains the muon g-2 anomaly while evading the experimental constraints. After establishing the viability of the singlet fermion to be a dark matter candidate, we calculate the positron excess produced from its annihilation to the light bosons which primarily decay to muons. Incorporating the Sommerfeld effect caused by the light mediator and an appropriate boost factor, we find that our proposed model can satisfactorily explain the positron fraction excess as well as the positron spectrum data reported by AMS-02 experiment.

hep-ph↗

Data-based Discovery of Governing Equations

Most common mechanistic models are traditionally presented in mathematical forms to explain a given physical phenomenon. Machine learning algorithms, on the other hand, provide a mechanism to map the input data to output without explicitly describing the underlying physical process that generated the data. We propose a Data-based Physics Discovery (DPD) framework for automatic discovery of governing equations from observed data. Without a prior definition of the model structure, first a free-form of the equation is discovered, and then calibrated and validated against the available data. In addition to the observed data, the DPD framework can utilize available prior physical models, and domain expert feedback. When prior models are available, the DPD framework can discover an additive or multiplicative correction term represented symbolically. The correction term can be a function of the existing input variable to the prior model, or a newly introduced variable. In case a prior model is not available, the DPD framework discovers a new data-based standalone model governing the observations. We demonstrate the performance of the proposed framework on a real-world application in the aerospace industry.

cs.LG↗

PRover: Proof Generation for Interpretable Reasoning over Rules

Recent work by Clark et al. (2020) shows that transformers can act as 'soft theorem provers' by answering questions over explicitly provided knowledge in natural language. In our work, we take a step closer to emulating formal theorem provers, by proposing PROVER, an interpretable transformer-based model that jointly answers binary questions over rule-bases and generates the corresponding proofs. Our model learns to predict nodes and edges corresponding to proof graphs in an efficient constrained training paradigm. During inference, a valid proof, satisfying a set of global constraints is generated. We conduct experiments on synthetic, hand-authored, and human-paraphrased rule-bases to show promising results for QA and proof generation, with strong generalization performance. First, PROVER generates proofs with an accuracy of 87%, while retaining or improving performance on the QA task, compared to RuleTakers (up to 6% improvement on zero-shot evaluation). Second, when trained on questions requiring lower depths of reasoning, it generalizes significantly better to higher depths (up to 15% improvement). Third, PROVER obtains near perfect QA accuracy of 98% using only 40% of the training data. However, generating proofs for questions requiring higher depths of reasoning becomes challenging, and the accuracy drops to 65% for 'depth 5', indicating significant scope for future work. Our code and models are publicly available at https://github.com/swarnaHub/PRover

cs.CL↗

Data-Informed Decomposition for Localized Uncertainty Quantification of Dynamical Systems

Industrial dynamical systems often exhibit multi-scale response due to material heterogeneities, operation conditions and complex environmental loadings. In such problems, it is the case that the smallest length-scale of the systems dynamics controls the numerical resolution required to effectively resolve the embedded physics. In practice however, high numerical resolutions is only required in a confined region of the system where fast dynamics or localized material variability are exhibited, whereas a coarser discretization can be sufficient in the rest majority of the system. To this end, a unified computational scheme with uniform spatio-temporal resolutions for uncertainty quantification can be very computationally demanding. Partitioning the complex dynamical system into smaller easier-to-solve problems based of the localized dynamics and material variability can reduce the overall computational cost. However, identifying the region of interest for high-resolution and intensive uncertainty quantification can be a problem dependent. The region of interest can be specified based on the localization features of the solution, user interest, and correlation length of the random material properties. For problems where a region of interest is not evident, Bayesian inference can provide a feasible solution. In this work, we employ a Bayesian framework to update our prior knowledge on the localized region of interest using measurements and system response. To address the computational cost of the Bayesian inference, we construct a Gaussian process surrogate for the forward model. Once, the localized region of interest is identified, we use polynomial chaos expansion to propagate the localization uncertainty. We demonstrate our framework through numerical experiments on a three-dimensional elastodynamic problem.

physics.comp-ph↗

Bayesian learning of orthogonal embeddings for multi-fidelity Gaussian Processes

We present a Bayesian approach to identify optimal transformations that map model input points to low dimensional latent variables. The "projection" mapping consists of an orthonormal matrix that is considered a priori unknown and needs to be inferred jointly with the GP parameters, conditioned on the available training data. The proposed Bayesian inference scheme relies on a two-step iterative algorithm that samples from the marginal posteriors of the GP parameters and the projection matrix respectively, both using Markov Chain Monte Carlo (MCMC) sampling. In order to take into account the orthogonality constraints imposed on the orthonormal projection matrix, a Geodesic Monte Carlo sampling algorithm is employed, that is suitable for exploiting probability measures on manifolds. We extend the proposed framework to multi-fidelity models using GPs including the scenarios of training multiple outputs together. We validate our framework on three synthetic problems with a known lower-dimensional subspace. The benefits of our proposed framework, are illustrated on the computationally challenging three-dimensional aerodynamic optimization of a last-stage blade for an industrial gas turbine, where we study the effect of an 85-dimensional airfoil shape parameterization on two output quantities of interest, specifically on the aerodynamic efficiency and the degree of reaction.

stat.ML↗