Searcharxiv⌕ Search

arXiv subjects

Sayan Ghosh

Publications and source records attributed to Sayan Ghosh.

At least 55 records · Page 3Linked to original sources

Heterogenous Multi-Source Data Fusion Through Input Mapping and Latent Variable Gaussian Process

Artificial intelligence and machine learning frameworks have served as computationally efficient mapping between inputs and outputs for engineering problems. These mappings have enabled optimization and analysis routines that have warranted superior designs, ingenious material systems and optimized manufacturing processes. A common occurrence in such modeling endeavors is the existence of multiple source of data, each differentiated by fidelity, operating conditions, experimental conditions, and more. Data fusion frameworks have opened the possibility of combining such differentiated sources into single unified models, enabling improved accuracy and knowledge transfer. However, these frameworks encounter limitations when the different sources are heterogeneous in nature, i.e., not sharing the same input parameter space. These heterogeneous input scenarios can occur when the domains differentiated by complexity, scale, and fidelity require different parametrizations. Towards addressing this void, a heterogeneous multi-source data fusion framework is proposed based on input mapping calibration (IMC) and latent variable Gaussian process (LVGP). In the first stage, the IMC algorithm is utilized to transform the heterogeneous input parameter spaces into a unified reference parameter space. In the second stage, a multi-source data fusion model enabled by LVGP is leveraged to build a single source-aware surrogate model on the transformed reference space. The proposed framework is demonstrated and analyzed on three engineering case studies (design of cantilever beam, design of ellipsoidal void and modeling properties of Ti6Al4V alloy). The results indicate that the proposed framework provides improved predictive accuracy over a single source model and transformed but source unaware model.

stat.ML↗

Pulse Shape Simulation and Discrimination using Machine-Learning Techniques

An essential metric for the quality of a particle-identification experiment is its statistical power to discriminate between signal and background. Pulse shape discrimination (PSD) is a basic method for this purpose in many nuclear, high-energy and rare-event search experiments where scintillation detectors are used. Conventional techniques exploit the difference between decay-times of the pulses from signal and background events or pulse signals caused by different types of radiation quanta to achieve good discrimination. However, such techniques are efficient only when the total light-emission is sufficient to get a proper pulse profile. This is only possible when adequate amount of energy is deposited from recoil of the electrons or the nuclei of the scintillator materials caused by the incident particle on the detector. But, rare-event search experiments like direct search for dark matter do not always satisfy these conditions. Hence, it becomes imperative to have a method that can deliver a very efficient discrimination in these scenarios. Neural network based machine-learning algorithms have been used for classification problems in many areas of physics especially in high-energy experiments and have given better results compared to conventional techniques. We present the results of our investigations of two network based methods \viz Dense Neural Network and Recurrent Neural Network, for pulse shape discrimination and compare the same with conventional methods.

physics.ins-det↗

Inhomogeneous Polarization Transformation Reveals PT-Transition in non-Hermitian Optical Beam Shift

Despite its non-Hermitian nature, the transverse optical beam shift exhibits both real eigenvalues and non-orthogonal eigenstates. To explore this unexpected similarity to typical PT (parity-time)-symmetric systems, we first categorize the entire parametric regime of optical beam shifts into Hermitian, PT-unbroken, and PT-broken phases. Besides experimentally unveiling the PT-broken regime, crucially, we illustrate that the observed PT-transition is rooted in the momentum-domain inhomogeneous polarization transformation of the beam. The correspondence with a typical non-Hermitian photonic system is further established. Our work not only resolves a longstanding fundamental issue in the field of optical beam shift but also puts forward the notion of novel non-Hermitian spin-orbit photonics: a new direction to study non-Hermitian physics through the optical beam shifts.

physics.optics↗

Emergent Quadrupolar Order in the Spin-$1/2$ Kitaev-Heisenberg Model

Motivated by the largely unexplored domain of multi-polar ordered spin states in the Kitaev-Heisenberg (KH) systems we investigate the ground state dynamics of the spin-$\frac{1}{2}$ KH model, focusing on quadrupolar (QP) order in 2-leg ladder and two-dimensional honeycomb lattice geometries. Employing exact diagonalization and density-matrix renormalization group methods, we analyze the QP order parameter and correlation functions. Our findings reveal a robust QP order across a wide range of the phase diagram, influenced by the interplay between Heisenberg and Kitaev interactions. Notably, we observe an enhancement of QP order near Kitaev quantum spin liquid (QSL) phases, despite the absence of long-range spin-spin correlations. This highlights a complex relationship between QP order and QSLs, offering new insights into quantum magnetism in low-dimensional systems. Our findings provide a rational explanation for the observed nonlinear magnetic susceptibility in $α$-RuCl$_3$.

cond-mat.str-el↗

Leveraging Multiple Teachers for Test-Time Adaptation of Language-Guided Classifiers

Recent approaches have explored language-guided classifiers capable of classifying examples from novel tasks when provided with task-specific natural language explanations, instructions or prompts (Sanh et al., 2022; R. Menon et al., 2022). While these classifiers can generalize in zero-shot settings, their task performance often varies substantially between different language explanations in unpredictable ways (Lu et al., 2022; Gonen et al., 2022). Also, current approaches fail to leverage unlabeled examples that may be available in many scenarios. Here, we introduce TALC, a framework that uses data programming to adapt a language-guided classifier for a new task during inference when provided with explanations from multiple teachers and unlabeled test examples. Our results show that TALC consistently outperforms a competitive baseline from prior work by an impressive 9.3% (relative improvement). Further, we demonstrate the robustness of TALC to variations in the quality and quantity of provided explanations, highlighting its potential in scenarios where learning from multiple teachers or a crowd is involved. Our code is available at: https://github.com/WeiKangda/TALC.git.

cs.CL↗

Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation Models

Generalized quantifiers (e.g., few, most) are used to indicate the proportions predicates are satisfied (for example, some apples are red). One way to interpret quantifier semantics is to explicitly bind these satisfactions with percentage scopes (e.g., 30%-40% of apples are red). This approach can be helpful for tasks like logic formalization and surface-form quantitative reasoning (Gordon and Schubert, 2010; Roy et al., 2015). However, it remains unclear if recent foundation models possess this ability, as they lack direct training signals. To explore this, we introduce QuRe, a crowd-sourced dataset of human-annotated generalized quantifiers in Wikipedia sentences featuring percentage-equipped predicates. We explore quantifier comprehension in language models using PRESQUE, a framework that combines natural language inference and the Rational Speech Acts framework. Experimental results on the HVD dataset and QuRe illustrate that PRESQUE, employing pragmatic reasoning, performs 20% better than a literal reasoning baseline when predicting quantifier percentage scopes, with no additional training required.

cs.AI↗

Beyond Labels: Empowering Human Annotators with Natural Language Explanations through a Novel Active-Learning Architecture

Real-world domain experts (e.g., doctors) rarely annotate only a decision label in their day-to-day workflow without providing explanations. Yet, existing low-resource learning techniques, such as Active Learning (AL), that aim to support human annotators mostly focus on the label while neglecting the natural language explanation of a data point. This work proposes a novel AL architecture to support experts' real-world need for label and explanation annotations in low-resource scenarios. Our AL architecture leverages an explanation-generation model to produce explanations guided by human explanations, a prediction model that utilizes generated explanations toward prediction faithfully, and a novel data diversity-based AL sampling strategy that benefits from the explanation annotations. Automated and human evaluations demonstrate the effectiveness of incorporating explanations into AL sampling and the improved human annotation efficiency and trustworthiness with our AL architecture. Additional ablation studies illustrate the potential of our AL architecture for transfer learning, generalizability, and integration with large language models (LLMs). While LLMs exhibit exceptional explanation-generation capabilities for relatively simple tasks, their effectiveness in complex real-world tasks warrants further in-depth study.

cs.CL↗

Signum phase mask differential microscopy

We propose and experimentally demonstrate a differential microscopy method to obtain simultaneous amplitude, phase, and quantitative polarization gradient imaging in a single experimental embodiment. A full-field optical spatial differentiator is achieved in a relatively simple setup by placing a glass cover slip as a Signum phase mask in the Fourier plane of a standard 4-f imaging system and accordingly named Signum phase mask differential microscopy. The longstanding requisite of polarized light to obtain the spatial differentiation of the field at the object plane is eliminated in our scheme and, hence, leads to the emergence of quantitative differential polarization contrast imaging by integrating polarization degree of freedom as an additional contrast agent in the framework of differential microscopy. Implementation of the proposed differential imaging scheme in high-resolution microscopy is experimentally demonstrated alongside its functionality for a broad wavelength range. Simultaneous acquisition of differential phase, amplitude, and polarization (anisotropy) gradient imaging in a rather elementary optical setup enables a low-cost multi-functional differential microscopy system that is anticipated to emerge as a revolutionary tool in label-free imaging and optical image processing.

physics.optics↗

Green Federated Learning

The rapid progress of AI is fueled by increasingly large and computationally intensive machine learning models and datasets. As a consequence, the amount of compute used in training state-of-the-art models is exponentially increasing (doubling every 10 months between 2015 and 2022), resulting in a large carbon footprint. Federated Learning (FL) - a collaborative machine learning technique for training a centralized model using data of decentralized entities - can also be resource-intensive and have a significant carbon footprint, particularly when deployed at scale. Unlike centralized AI that can reliably tap into renewables at strategically placed data centers, cross-device FL may leverage as many as hundreds of millions of globally distributed end-user devices with diverse energy sources. Green AI is a novel and important research area where carbon footprint is regarded as an evaluation criterion for AI, alongside accuracy, convergence speed, and other metrics. In this paper, we propose the concept of Green FL, which involves optimizing FL parameters and making design choices to minimize carbon emissions consistent with competitive performance and training time. The contributions of this work are two-fold. First, we adopt a data-driven approach to quantify the carbon emissions of FL by directly measuring real-world at-scale FL tasks running on millions of phones. Second, we present challenges, guidelines, and lessons learned from studying the trade-off between energy efficiency, performance, and time-to-train in a production FL system. Our findings offer valuable insights into how FL can reduce its carbon footprint, and they provide a foundation for future research in the area of Green AI.

cs.LG↗

Quantized topological energy pumping and Weyl points in Floquet synthetic dimensions with a driven-dissipative photonic molecule

Topological effects manifest in a wide range of physical systems, such as solid crystals, acoustic waves, photonic materials and cold atoms. These effects are characterized by `topological invariants' which are typically integer-valued, and lead to robust quantized channels of transport in space, time, and other degrees of freedom. The temporal channel, in particular, allows one to achieve higher-dimensional topological effects, by driving the system with multiple incommensurate frequencies. However, dissipation is generally detrimental to such topological effects, particularly when the systems consist of quantum spins or qubits. Here we introduce a photonic molecule subjected to multiple RF/optical drives and dissipation as a promising candidate system to observe quantized transport along Floquet synthetic dimensions. Topological energy pumping in the incommensurately modulated photonic molecule is enhanced by the driven-dissipative nature of our platform. Furthermore, we provide a path to realizing Weyl points and measuring the Berry curvature emanating from these reciprocal-space ($k$-space) magnetic monopoles, illustrating the capabilities for higher-dimensional topological Hamiltonian simulation in this platform. Our approach enables direct $k$-space engineering of a wide variety of Hamiltonians using modulation bandwidths that are well below the free-spectral range (FSR) of integrated photonic cavities.

physics.optics↗

Bridging Nations: Quantifying the Role of Multilinguals in Communication on Social Media

Social media enables the rapid spread of many kinds of information, from memes to social movements. However, little is known about how information crosses linguistic boundaries. We apply causal inference techniques on the European Twitter network to quantify multilingual users' structural role and communication influence in cross-lingual information exchange. Overall, multilinguals play an essential role; posting in multiple languages increases betweenness centrality by 13%, and having a multilingual network neighbor increases monolinguals' odds of sharing domains and hashtags from another language 16-fold and 4-fold, respectively. We further show that multilinguals have a greater impact on diffusing information less accessible to their monolingual compatriots, such as information from far-away countries and content about regional politics, nascent social movements, and job opportunities. By highlighting information exchange across borders, this work sheds light on a crucial component of how information and ideas spread around the world.

cs.SI↗

When does the student surpass the teacher? Federated Semi-supervised Learning with Teacher-Student EMA

Semi-Supervised Learning (SSL) has received extensive attention in the domain of computer vision, leading to development of promising approaches such as FixMatch. In scenarios where training data is decentralized and resides on client devices, SSL must be integrated with privacy-aware training techniques such as Federated Learning. We consider the problem of federated image classification and study the performance and privacy challenges with existing federated SSL (FSSL) approaches. Firstly, we note that even state-of-the-art FSSL algorithms can trivially compromise client privacy and other real-world constraints such as client statelessness and communication cost. Secondly, we observe that it is challenging to integrate EMA (Exponential Moving Average) updates into the federated setting, which comes at a trade-off between performance and communication cost. We propose a novel approach FedSwitch, that improves privacy as well as generalization performance through Exponential Moving Average (EMA) updates. FedSwitch utilizes a federated semi-supervised teacher-student EMA framework with two features - local teacher adaptation and adaptive switching between teacher and student for pseudo-label generation. Our proposed approach outperforms the state-of-the-art on federated image classification, can be adapted to real-world constraints, and achieves good generalization performance with minimal communication cost overhead.

cs.LG↗

Pruning Compact ConvNets for Efficient Inference

Neural network pruning is frequently used to compress over-parameterized networks by large amounts, while incurring only marginal drops in generalization performance. However, the impact of pruning on networks that have been highly optimized for efficient inference has not received the same level of attention. In this paper, we analyze the effect of pruning for computer vision, and study state-of-the-art ConvNets, such as the FBNetV3 family of models. We show that model pruning approaches can be used to further optimize networks trained through NAS (Neural Architecture Search). The resulting family of pruned models can consistently obtain better performance than existing FBNetV3 models at the same level of computation, and thus provide state-of-the-art results when trading off between computational complexity and generalization performance on the ImageNet benchmark. In addition to better generalization performance, we also demonstrate that when limited computation resources are available, pruning FBNetV3 models incur only a fraction of GPU-hours involved in running a full-scale NAS.

cs.CV↗

LaSQuE: Improved Zero-Shot Classification from Explanations Through Quantifier Modeling and Curriculum Learning

A hallmark of human intelligence is the ability to learn new concepts purely from language. Several recent approaches have explored training machine learning models via natural language supervision. However, these approaches fall short in leveraging linguistic quantifiers (such as 'always' or 'rarely') and mimicking humans in compositionally learning complex tasks. Here, we present LaSQuE, a method that can learn zero-shot classifiers from language explanations by using three new strategies - (1) modeling the semantics of linguistic quantifiers in explanations (including exploiting ordinal strength relationships, such as 'always' > 'likely'), (2) aggregating information from multiple explanations using an attention-based mechanism, and (3) model training via curriculum learning. With these strategies, LaSQuE outperforms prior work, showing an absolute gain of up to 7% in generalizing to unseen real-world classification tasks.

cs.CL↗

A Comprehensive Review of Digital Twin -- Part 1: Modeling and Twinning Enabling Technologies

As an emerging technology in the era of Industry 4.0, digital twin is gaining unprecedented attention because of its promise to further optimize process design, quality control, health monitoring, decision and policy making, and more, by comprehensively modeling the physical world as a group of interconnected digital models. In a two-part series of papers, we examine the fundamental role of different modeling techniques, twinning enabling technologies, and uncertainty quantification and optimization methods commonly used in digital twins. This first paper presents a thorough literature review of digital twin trends across many disciplines currently pursuing this area of research. Then, digital twin modeling and twinning enabling technologies are further analyzed by classifying them into two main categories: physical-to-virtual, and virtual-to-physical, based on the direction in which data flows. Finally, this paper provides perspectives on the trajectory of digital twin technology over the next decade, and introduces a few emerging areas of research which will likely be of great use in future digital twin research. In part two of this review, the role of uncertainty quantification and optimization are discussed, a battery digital twin is demonstrated, and more perspectives on the future of digital twin are shared.

cs.CE↗

A Comprehensive Review of Digital Twin -- Part 2: Roles of Uncertainty Quantification and Optimization, a Battery Digital Twin, and Perspectives

As an emerging technology in the era of Industry 4.0, digital twin is gaining unprecedented attention because of its promise to further optimize process design, quality control, health monitoring, decision and policy making, and more, by comprehensively modeling the physical world as a group of interconnected digital models. In a two-part series of papers, we examine the fundamental role of different modeling techniques, twinning enabling technologies, and uncertainty quantification and optimization methods commonly used in digital twins. This second paper presents a literature review of key enabling technologies of digital twins, with an emphasis on uncertainty quantification, optimization methods, open source datasets and tools, major findings, challenges, and future directions. Discussions focus on current methods of uncertainty quantification and optimization and how they are applied in different dimensions of a digital twin. Additionally, this paper presents a case study where a battery digital twin is constructed and tested to illustrate some of the modeling and twinning methods reviewed in this two-part review. Code and preprocessed data for generating all the results and figures presented in the case study are available on GitHub.

cs.LG↗

Opacus: User-Friendly Differential Privacy Library in PyTorch

We introduce Opacus, a free, open-source PyTorch library for training deep learning models with differential privacy (hosted at opacus.ai). Opacus is designed for simplicity, flexibility, and speed. It provides a simple and user-friendly API, and enables machine learning practitioners to make a training pipeline private by adding as little as two lines to their code. It supports a wide variety of layers, including multi-head attention, convolution, LSTM, GRU (and generic RNN), and embedding, right out of the box and provides the means for supporting other user-defined layers. Opacus computes batched per-sample gradients, providing higher efficiency compared to the traditional "micro batch" approach. In this paper we present Opacus, detail the principles that drove its implementation and unique features, and benchmark it against other frameworks for training models with differential privacy as well as standard PyTorch.

cs.LG↗

Inelastic charged current interaction of supernova neutrinos in two-phase liquid xenon dark matter detectors

It has been known that neutrinos from supernova (SN) bursts can give rise to nuclear recoil (NR) signals arising from coherent elastic neutrino-nucleus scattering (CE$ν$NS) interaction, a neutral current (NC) process, of the neutrinos with xenon nuclei in future large (multi-ton scale) liquid xenon (LXe) detectors employed for dark matter search, depending on the SN progenitor mass and distance to the SN. In this paper, we show that the same detectors will also be sensitive to inelastic charged current (CC) interactions of the SN electron neutrinos ($ν_e$CC) with the xenon nuclei. Such interactions, while creating an electron in the final state, also leave the post-interaction target nucleus in an excited state, the subsequent deexcitation of which produces, among other particles, gamma rays and neutrons. The electron and deexcitation gamma rays will give ``electron recoil" (ER) type signals, while the deexcitation neutrons produce, through their multiple scattering on the xenon nuclei, further xenon nuclear recoils that will also give NR signals (in addition to those produced through the CE$ν$NS interactions). We discuss the observable scintillation and ionization signals associated with SN neutrino induced CE$ν$NS and $ν_e$CC events in a generic LXe detector and argue that upcoming sufficiently large LXe detectors should be able to detect both these types of events due to neutrinos from reasonably close by SN bursts. We also note that since the total CC induced ER and NR signals receive contributions predominantly from $ν_e$CC interactions while the CE$ν$NS contribution comes from NC interactions of {\emph all the six species of neutrinos}, identification of the $ν_e$CC and CE$ν$NS origin events may offer the possibility of extracting useful information about the distribution of the total SN explosion energy going into different neutrino flavors.

hep-ph↗