Searcharxiv⌕ Search

arXiv subjects

Souradeep Chattopadhyay

Publications and source records attributed to Souradeep Chattopadhyay.

5 recordsLinked to original sources

Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models

In the critical task of making generative models trustworthy and robust, methods for Uncertainty Quantification (UQ) have begun to show encouraging potential. However, many of these methods rely on rigid heuristics that fail to generalize across tasks and modalities. Here, we propose a novel framework for UQ that is highly flexible and approaches or surpasses the performance of prior heuristic methods. We introduce Directional Concentration Uncertainty (DCU), a novel statistical procedure for quantifying the concentration of embeddings based on the von Mises-Fisher (vMF) distribution. Our method captures uncertainty by measuring the geometric dispersion of multiple generated outputs from a language model using continuous embeddings of the generated outputs without any task specific heuristics. In our experiments, we show that DCU matches or exceeds calibration levels of prior works like semantic entropy (Kuhn et al., 2023) and also generalizes well to more complex tasks in multi-modal domains. We present a framework for the wider potential of DCU and its implications for integration into UQ for multi-modal and agentic frameworks.

cs.LG↗

Predicting Mild Cognitive Impairment Using Naturalistic Driving and Trip Destination Modeling

Understanding the relationship between mild cognitive impairment (MCI) and driving behavior is essential for enhancing road safety, particularly among older adults. This study introduces a novel approach by incorporating specific trip destinations-such as home, work, medical appointments, social activities, and errands-using geohashing to analyze the driving habits of older drivers in Nebraska. We employed a two-fold methodology that combines data visualization with advanced machine learning models, including C5.0, Random Forest, and Support Vector Machines, to assess the effectiveness of these location-based variables in predicting cognitive impairment. Notably, the C5.0 model showed a robust and stable performance, achieving a median recall of 0.68, which indicates that our methodology accurately identifies cognitive impairment in drivers 68\% of the time. This emphasizes our model's capacity to reduce false negatives, a crucial factor given the profound implications of failing to identify impaired drivers. Our findings underscore the innovative use of life-space variables in understanding and predicting cognitive decline, offering avenues for early intervention and tailored support for affected individuals.

cs.LG↗

Multi-layered characterization of hot stellar systems with confidence

Understanding the physical and evolutionary properties of Hot Stellar Systems (HSS) is a major challenge in astronomy. We studied the dataset on 13456 HSS of Misgeld and Hilker (2011) that includes 12763 candidate globular clusters using stellar mass ($M_s$), effective radius ($R_e$) and mass-to-luminosity ratio ($M_s/L_ν$), and found multi-layered homogeneous grouping among these stellar systems. Our methods elicited eight homogeneous ellipsoidal groups at the finest sub-group level. Some of these groups have high overlap and were merged through a multi-phased syncytial algorithm motivated from Almodóvar-Rivera and Maitra (2020). Five groups were merged in the first phase, resulting in three complex-structured groups. Our algorithm determined further complex structure and permitted another merging phase, revealing two complex-structured groups at the highest level. A nonparametric bootstrap procedure was also used to estimate the confidence of each of our group assignments. These assignments generally had high confidence in classification, indicating great degree of certainty of the HSS assignments into our complex-structured groups. The physical and kinematic properties of the two groups were assessed in terms of $M_s$, $R_e$, surface density and $M_s/L_ν$. The first group consisted of older, smaller and less bright HSS while the second group consisted of brighter and younger HSS. Our analysis provides novel insight into the physical and evolutionary properties of HSS and also helps understand physical and evolutionary properties of candidate globular clusters. Further, the candidate globular clusters (GCs) are seen to have very high chance of really being GCs rather than dwarfs or dwarf ellipticals that are also indicated to be quite distinct from each other.

astro-ph.GA↗

Multivariate $t$-Mixtures-Model-based Cluster Analysis of BATSE Catalog Establishes Importance of All Observed Parameters, Confirms Five Distinct Ellipsoidal Sub-populations of Gamma Ray Bursts

Determining the kinds of gamma-ray bursts (GRBs) has been of interest to astronomers for many years. We analyzed 1599 GRBs from the Burst and Transient Source Experiment (BATSE) 4Br catalogue using $t$-mixtures-model-based clustering on all nine observed parameters ($T_{50}$, $T_{90}$, $F_1$, $F_2$, $F_3$, $F_4$, $P_{64}$, $P_{256}$, $P_{1024}$) and found evidence of five types of GRBs. Our results further refine the findings of Chattopadhyay and Maitra (2017) by providing groups that are more distinct. Using the Mukherjee et al. (1998) classification scheme, also used by Chattopadhyay and Maitra (2017), of duration, total fluence ($F_t = F_1 + F_2 + F_3 + F_4$)) and spectrum (using Hardness Ratio $H_{321} = F_3/(F_1 + F_2)$) our five groups are classified as long-intermediate-intermediate, short-faint-intermediate, short-faint-soft, long-bright-hard, and long-intermediate-hard. We also classify 374 GRBs in the BATSE catalogue that have incomplete information in some of the observed variables (mainly the four time integrated fluences $F_1$, $F_2$, $F_3$ and $F_4$) to the five groups obtained, using the 1599 GRBs having complete information in all the observed variables. Our classification scheme puts 138 GRBs in the first group, 52 GRBs in the second group, 33 GRBs in the third group, 127 GRBs in the fourth group and 24 GRBs in the fifth group.

astro-ph.HE↗

Gaussian-Mixture-Model-based Cluster Analysis Finds Five Kinds of Gamma Ray Bursts in the BATSE Catalog

Clustering methods are an important tool to enumerate and describe the different coherent kinds of Gamma Ray Bursts (GRBs). But their performance can be affected by a number of factors such as the choice of clustering algorithm and inherent associated assumptions, the inclusion of variables in clustering, nature of initialization methods used or the iterative algorithm or the criterion used to judge the optimal number of groups supported by the data. We analyzed GRBs from the BATSE 4Br catalog using $k$-means and Gaussian Mixture Models-based clustering methods and found that after accounting for all the above factors, all six variables -- different subsets of which have been used in the literature -- and that are, namely, the flux duration variables ($T_{50}$, $T_{90}$), the peak flux ($P_{256}$) measured in 256-millisecond bins, the total fluence ($F_t$) and the spectral hardness ratios ($H_{32}$ and $H_{321}$) contain information on clustering. Further, our analysis found evidence of five different kinds of GRBs and that these groups have different kinds of dispersions in terms of shape, size and orientation. In terms of duration, fluence and spectrum, the five types of GRBs were characterized as intermediate/faint/intermediate, long/intermediate/soft, intermediate/intermediate/intermediate, short/faint/hard and long/bright/intermediate.

astro-ph.HE↗