Searcharxiv⌕ Search

arXiv subjects

Joydeep Ghosh

Publications and source records attributed to Joydeep Ghosh.

At least 37 records · Page 2Linked to original sources

Identification and Characterization of a New Disruption Regime in ADITYA-U Tokamak

Disruptions continue to pose a significant challenge to the stable operation and future design of tokamak reactors. A comprehensive statistical investigation carried out on the ADITYA-U tokamak has led to the observation and characterization of a novel disruption regime. In contrast to the conventional Locked Mode Disruption (LMD), the newly identified disruption exhibits a distinctive two-phase evolution: an initial phase characterized by a steady rise in mode frequency with a nonlinearly saturated amplitude, followed by a sudden frequency collapse accompanied by a pronounced increase in amplitude. This behaviour signifies the onset of the precursor phase on a significantly shorter timescale. Clear empirical thresholds have been identified to distinguish this disruption type from conventional LMD events, including edge safety factor, current decay coefficient, current quench (CQ) time, and CQ rate. The newly identified disruption regime is predominantly governed by the (m/n = 2/1) drift-tearing mode (DTM), which, in contrast to typical disruptions in the ADITYA-U tokamak that involve both m/n = 2/1 and 3/1 modes, consistently manifests as the sole dominant instability. Initiated by core temperature hollowing, the growth of this mode is significantly enhanced by a synergistic interplay between a strongly localized pressure gradient and the pronounced steepening of the current density profile in the vicinity of the mode rational surface.

physics.plasm-ph↗

Effect of convective transport in edge/SOL plasmas of ADITYA-U tokamak

The 2-D edge plasma fluid transport code, UEDGE has been used to simulate the edge region of circular limiter plasmas of ADITYA-U for modelling the measured electron density profile. The limiter geometry of ADITYA-U has been introduced in the UEDGE code, which is primarily developed and used for divertor configuration. The computational mesh defining the limiter geometry is generated by a routine developed in-house, and has successfully been integrated with the UEDGE code to simulate the edge plasma parameters of ADITYA-U. The radial profiles of edge and scrape-off layer (SOL) electron density, ne and temperature are obtained from the simulations and used to model the measured ne profile using Langmuir probe array. It has been found that a convective velocity, vconv. is definitely needed in addition to the constant perpendicular diffusion coefficient, D throughout the edge and SOL regions to model the edge ne profile. The obtained vconv. is inward and radially constant with a value of 1.5 m/s and the radially constant D is ~ 0.2 m 2/s. The value of D ~ 0.2 m2/s is found to be much less than fluctuation induced diffusivities and lies in-between the neoclassical diffusivity and Bohm diffusivity estimated in the edge-SOL region of ADITYA-U tokamak. Furthermore, the transport of radial electron heat flux is found to be maximizing near the limiter tip location in the poloidal plane.

physics.plasm-ph↗

A multi-purpose reciprocating probe drive system for studying the effect of gas-puffs on edge plasma dynamics in the ADITYA-U tokamak

This article reports the development of a versatile high-speed reciprocating drive system (HRDS) with interchangeable probe heads to characterize the edge plasma region of ADITYA-U tokamak. This reciprocating probe drive system consisting of Langmuir and magnetic probe heads, is designed, fabricated, installed, and operated for studying the extent of fuel/impurity gas propagation and its influence on plasma dynamics in the far-edge region inside the last closed magnetic flux surface (LCFS). The HRDS is driven by a highly accurate, easy-to-control, dynamic, brushless, permanently excited synchronous servo motor operated by a PXI-commanded controller. The system is remotely operated and allows for precise control of the speed, acceleration, and distance traveled of the probe head on a shot-to-shot basis, facilitating seamless control of operations according to experimental requirements. Using this system, consisting of a linear array of Langmuir probes, measurements of plasma density, temperature, potential, and their fluctuations revealed that the fuel gas-puff impact these mean and fluctuating parameters up to three to four cm inside the LCFS. Attaching an array of magnetic probes to this system led to measurements of magnetic fluctuations inside the LCFS. The HRDS system is fully operational and serves as an important diagnostic tool for ADITYA-U tokamak.

physics.plasm-ph↗

MHD activity induced coherent mode excitation in the edge plasma region of ADITYA-U Tokamak

In this paper, we report the excitation of coherent density and potential fluctuations induced by magnetohydrodynamic (MHD) activity in the edge plasma region of ADITYA-U Tokamak. When the amplitude of the MHD mode, mainly the m/n = 2/1, increases beyond a threshold value of 0.3-0.4 %, coherent oscillations in the density and potential fluctuations are observed having the same frequency as that of the MHD mode. The mode numbers of these MHD induced density and potential fluctuations are obtained by Langmuir probes placed at different radial, poloidal, and toroidal locations in the edge plasma region. Detailed analyses of these Langmuir probe measurements reveal that the coherent mode in edge potential fluctuation has a mode structure of m/n = 2/1 whereas the edge density fluctuation has an m/n = 1/1 structure. It is further observed that beyond the threshold, the coupled power fraction scales almost linearly with the magnitude of magnetic fluctuations. Furthermore, the rise rates of the coupled power fraction for coherent modes in density and potential fluctuations are also found to be dependent on the growth rate of magnetic fluctuations. The disparate mode structures of the excited modes in density and plasma potential fluctuations suggest that the underlying mechanism for their existence is most likely due to the excitation of the global high-frequency branch of zonal flows occurring through the coupling of even harmonics of potential to the odd harmonics of pressure due to 1/R dependence of the toroidal magnetic field.

physics.plasm-ph↗

Novel Node Category Detection Under Subpopulation Shift

In real-world graph data, distribution shifts can manifest in various ways, such as the emergence of new categories and changes in the relative proportions of existing categories. It is often important to detect nodes of novel categories under such distribution shifts for safety or insight discovery purposes. We introduce a new approach, Recall-Constrained Optimization with Selective Link Prediction (RECO-SLIP), to detect nodes belonging to novel categories in attributed graphs under subpopulation shifts. By integrating a recall-constrained learning framework with a sample-efficient link prediction mechanism, RECO-SLIP addresses the dual challenges of resilience against subpopulation shifts and the effective exploitation of graph structure. Our extensive empirical evaluation across multiple graph datasets demonstrates the superior performance of RECO-SLIP over existing methods. The experimental code is available at https://github.com/hsinghuan/novel-node-category-detection.

cs.LG↗

Federated Learning for Estimating Heterogeneous Treatment Effects

Machine learning methods for estimating heterogeneous treatment effects (HTE) facilitate large-scale personalized decision-making across various domains such as healthcare, policy making, education, and more. Current machine learning approaches for HTE require access to substantial amounts of data per treatment, and the high costs associated with interventions makes centrally collecting so much data for each intervention a formidable challenge. To overcome this obstacle, in this work, we propose a novel framework for collaborative learning of HTE estimators across institutions via Federated Learning. We show that even under a diversity of interventions and subject populations across clients, one can jointly learn a common feature representation, while concurrently and privately learning the specific predictive functions for outcomes under distinct interventions across institutions. Our framework and the associated algorithm are based on this insight, and leverage tabular transformers to map multiple input data to feature representations which are then used for outcome prediction via multi-task learning. We also propose a novel way of federated training of personalised transformers that can work with heterogeneous input feature spaces. Experimental results on real-world clinical trial data demonstrate the effectiveness of our method.

cs.LG↗

Achieving Fairness Across Local and Global Models in Federated Learning

Achieving fairness across diverse clients in Federated Learning (FL) remains a significant challenge due to the heterogeneity of the data and the inaccessibility of sensitive attributes from clients' private datasets. This study addresses this issue by introducing \texttt{EquiFL}, a novel approach designed to enhance both local and global fairness in federated learning environments. \texttt{EquiFL} incorporates a fairness term into the local optimization objective, effectively balancing local performance and fairness. The proposed coordination mechanism also prevents bias from propagating across clients during the collaboration phase. Through extensive experiments across multiple benchmarks, we demonstrate that \texttt{EquiFL} not only strikes a better balance between accuracy and fairness locally at each client but also achieves global fairness. The results also indicate that \texttt{EquiFL} ensures uniform performance distribution among clients, thus contributing to performance fairness. Furthermore, we showcase the benefits of \texttt{EquiFL} in a real-world distributed dataset from a healthcare application, specifically in predicting the effects of treatments on patients across various hospital locations.

cs.LG↗

SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors

Popular parameter-efficient fine-tuning (PEFT) methods, such as LoRA and its variants, freeze pre-trained model weights \(W\) and inject learnable matrices \(ΔW\). These \(ΔW\) matrices are structured for efficient parameterization, often using techniques like low-rank approximations or scaling vectors. However, these methods typically show a performance gap compared to full fine-tuning. Although recent PEFT methods have narrowed this gap, they do so at the cost of additional learnable parameters. We propose SVFT, a simple approach that fundamentally differs from existing methods: the structure imposed on \(ΔW\) depends on the specific weight matrix \(W\). Specifically, SVFT updates \(W\) as a sparse combination of outer products of its singular vectors, training only the coefficients (scales) of these sparse combinations. This approach allows fine-grained control over expressivity through the number of coefficients. Extensive experiments on language and vision benchmarks show that SVFT recovers up to 96% of full fine-tuning performance while training only 0.006 to 0.25% of parameters, outperforming existing methods that only recover up to 85% performance using 0.03 to 0.8% of the trainable parameter budget.

cs.LG↗

Exploring Explainability in Video Action Recognition

Image Classification and Video Action Recognition are perhaps the two most foundational tasks in computer vision. Consequently, explaining the inner workings of trained deep neural networks is of prime importance. While numerous efforts focus on explaining the decisions of trained deep neural networks in image classification, exploration in the domain of its temporal version, video action recognition, has been scant. In this work, we take a deeper look at this problem. We begin by revisiting Grad-CAM, one of the popular feature attribution methods for Image Classification, and its extension to Video Action Recognition tasks and examine the method's limitations. To address these, we introduce Video-TCAV, by building on TCAV for Image Classification tasks, which aims to quantify the importance of specific concepts in the decision-making process of Video Action Recognition models. As the scalable generation of concepts is still an open problem, we propose a machine-assisted approach to generate spatial and spatiotemporal concepts relevant to Video Action Recognition for testing Video-TCAV. We then establish the importance of temporally-varying concepts by demonstrating the superiority of dynamic spatiotemporal concepts over trivial spatial concepts. In conclusion, we introduce a framework for investigating hypotheses in action recognition and quantitatively testing them, thus advancing research in the explainability of deep neural networks used in video action recognition.

cs.CV↗

Uncovering Misattributed Suicide Causes through Annotation Inconsistency Detection in Death Investigation Notes

Data accuracy is essential for scientific research and policy development. The National Violent Death Reporting System (NVDRS) data is widely used for discovering the patterns and causes of death. Recent studies suggested the annotation inconsistencies within the NVDRS and the potential impact on erroneous suicide-cause attributions. We present an empirical Natural Language Processing (NLP) approach to detect annotation inconsistencies and adopt a cross-validation-like paradigm to identify problematic instances. We analyzed 267,804 suicide death incidents between 2003 and 2020 from the NVDRS. Our results showed that incorporating the target state's data into training the suicide-crisis classifier brought an increase of 5.4% to the F-1 score on the target state's test set and a decrease of 1.1% on other states' test set. To conclude, we demonstrated the annotation inconsistencies in NVDRS's death investigation notes, identified problematic instances, evaluated the effectiveness of correcting problematic instances, and eventually proposed an NLP improvement solution.

cs.CL↗

Designing Robust Transformers using Robust Kernel Density Estimation

Recent advances in Transformer architectures have empowered their empirical success in a variety of tasks across different domains. However, existing works mainly focus on predictive accuracy and computational cost, without considering other practical issues, such as robustness to contaminated samples. Recent work by Nguyen et al., (2022) has shown that the self-attention mechanism, which is the center of the Transformer architecture, can be viewed as a non-parametric estimator based on kernel density estimation (KDE). This motivates us to leverage a set of robust kernel density estimation methods for alleviating the issue of data contamination. Specifically, we introduce a series of self-attention mechanisms that can be incorporated into different Transformer architectures and discuss the special properties of each method. We then perform extensive empirical studies on language modeling and image classification tasks. Our methods demonstrate robust performance in multiple scenarios while maintaining competitive results on clean datasets.

cs.LG↗

Privacy Preserving Bayesian Federated Learning in Heterogeneous Settings

In several practical applications of federated learning (FL), the clients are highly heterogeneous in terms of both their data and compute resources, and therefore enforcing the same model architecture for each client is very limiting. Moreover, the need for uncertainty quantification and data privacy constraints are often particularly amplified for clients that have limited local data. This paper presents a unified FL framework to simultaneously address all these constraints and concerns, based on training customized local Bayesian models that learn well even in the absence of large local datasets. A Bayesian framework provides a natural way of incorporating supervision in the form of prior distributions. We use priors in the functional (output) space of the networks to facilitate collaboration across heterogeneous clients. Moreover, formal differential privacy guarantees are provided for this framework. Experiments on standard FL datasets demonstrate that our approach outperforms strong baselines in both homogeneous and heterogeneous settings and under strict privacy constraints, while also providing characterizations of model uncertainties.

cs.LG↗

Behaviour of Ion Acoustic Soliton in a two-electron temperature plasmas of Multi-pole line cusp Plasma Device (MPD)

This article presents the experimental observations and characterization of Ion Acoustic Soliton (IAS) in a unique Multi-pole line cusp Plasma Device (MPD) device in which the magnitude of the pole-cusp magnetic field can be varied. And by varying the magnitude of the pole-cusp magnetic field, the proportions of two-electron-temperature components in the filament-produced plasmas of MPD can be varied. The solitons are experimentally characterized by measuring their amplitude-width relation and Mach numbers. The nature of the solitons is further established by making two counter-propagating solitons interact with each other. Later, the effect of the two-temperature electron population on soliton amplitude and width is studied by varying the magnitude of the pole cusp-magnetic field. It has been observed that different proportions of two-electron-temperature significantly influence the propagation of IAS. The amplitude of the soliton has been found to be following inversely with the effective electron temperature (Teff)

physics.plasm-ph↗

Split Localized Conformal Prediction

Conformal prediction is a simple and powerful tool that can quantify uncertainty without any distributional assumptions. Many existing methods only address the average coverage guarantee, which is not ideal compared to the stronger conditional coverage guarantee. Existing methods of approximating conditional coverage require additional models or time effort, which makes them not easy to scale. In this paper, we propose a modified non-conformity score by leveraging the local approximation of the conditional distribution using kernel density estimation. The modified score inherits the spirit of split conformal methods, which is simple and efficient and can scale to high dimensional settings. We also proposed a unified framework that brings together our method and several state-of-the-art. We perform extensive empirical evaluations: results measured by both average and conditional coverage confirm the advantage of our method.

stat.ML↗

Dynamic Combination of Heterogeneous Models for Hierarchical Time Series

We introduce a framework to dynamically combine heterogeneous models called \texttt{DYCHEM}, which forecasts a set of time series that are related through an aggregation hierarchy. Different types of forecasting models can be employed as individual ``experts'' so that each model is tailored to the nature of the corresponding time series. \texttt{DYCHEM} learns hierarchical structures during the training stage to help generalize better across all the time series being modeled and also mitigates coherency issues that arise due to constraints imposed by the hierarchy. To improve the reliability of forecasts, we construct quantile estimations based on the point forecasts obtained from combined heterogeneous models. The resulting quantile forecasts are coherent and independent of the choice of forecasting models. We conduct a comprehensive evaluation of both point and quantile forecasts for hierarchical time series (HTS), including public data and user records from a large financial software company. In general, our method is robust, adaptive to datasets with different properties, and highly configurable and efficient for large-scale forecasting pipelines.

cs.LG↗

Intermediate Entity-based Sparse Interpretable Representation Learning

Interpretable entity representations (IERs) are sparse embeddings that are "human-readable" in that dimensions correspond to fine-grained entity types and values are predicted probabilities that a given entity is of the corresponding type. These methods perform well in zero-shot and low supervision settings. Compared to standard dense neural embeddings, such interpretable representations may permit analysis and debugging. However, while fine-tuning sparse, interpretable representations improves accuracy on downstream tasks, it destroys the semantics of the dimensions which were enforced in pre-training. Can we maintain the interpretable semantics afforded by IERs while improving predictive performance on downstream tasks? Toward this end, we propose Intermediate enTity-based Sparse Interpretable Representation Learning (ItsIRL). ItsIRL realizes improved performance over prior IERs on biomedical tasks, while maintaining "interpretability" generally and their ability to support model debugging specifically. The latter is enabled in part by the ability to perform "counterfactual" fine-grained entity type manipulation, which we explore in this work. Finally, we propose a method to construct entity type based class prototypes for revealing global semantic properties of classes learned by our model.

cs.CL↗

FASTER-CE: Fast, Sparse, Transparent, and Robust Counterfactual Explanations

Counterfactual explanations have substantially increased in popularity in the past few years as a useful human-centric way of understanding individual black-box model predictions. While several properties desired of high-quality counterfactuals have been identified in the literature, three crucial concerns: the speed of explanation generation, robustness/sensitivity and succinctness of explanations (sparsity) have been relatively unexplored. In this paper, we present FASTER-CE: a novel set of algorithms to generate fast, sparse, and robust counterfactual explanations. The key idea is to efficiently find promising search directions for counterfactuals in a latent space that is specified via an autoencoder. These directions are determined based on gradients with respect to each of the original input features as well as of the target, as estimated in the latent space. The ability to quickly examine combinations of the most promising gradient directions as well as to incorporate additional user-defined constraints allows us to generate multiple counterfactual explanations that are sparse, realistic, and robust to input manipulations. Through experiments on three datasets of varied complexities, we show that FASTER-CE is not only much faster than other state of the art methods for generating multiple explanations but also is significantly superior when considering a larger set of desirable (and often conflicting) properties. Specifically we present results across multiple performance metrics: sparsity, proximity, validity, speed of generation, and the robustness of explanations, to highlight the capabilities of the FASTER-CE family.

cs.LG↗

FEAMOE: Fair, Explainable and Adaptive Mixture of Experts

Three key properties that are desired of trustworthy machine learning models deployed in high-stakes environments are fairness, explainability, and an ability to account for various kinds of "drift". While drifts in model accuracy, for example due to covariate shift, have been widely investigated, drifts in fairness metrics over time remain largely unexplored. In this paper, we propose FEAMOE, a novel "mixture-of-experts" inspired framework aimed at learning fairer, more explainable/interpretable models that can also rapidly adjust to drifts in both the accuracy and the fairness of a classifier. We illustrate our framework for three popular fairness measures and demonstrate how drift can be handled with respect to these fairness constraints. Experiments on multiple datasets show that our framework as applied to a mixture of linear experts is able to perform comparably to neural networks in terms of accuracy while producing fairer models. We then use the large-scale HMDA dataset and show that while various models trained on HMDA demonstrate drift with respect to both accuracy and fairness, FEAMOE can ably handle these drifts with respect to all the considered fairness measures and maintain model accuracy as well. We also prove that the proposed framework allows for producing fast Shapley value explanations, which makes computationally efficient feature attribution based explanations of model decisions readily available via FEAMOE.

cs.LG↗