SearcharxivSearch

arXiv subjects

Amit Kumar Das

Publications and source records attributed to Amit Kumar Das.

10 recordsLinked to original sources

What Catches the Eye? A Conjoint Study of Infographic Design Preferences

Infographic designers balance many choices at once: chart type, color, and whether to add a benchmark or a scale. Past work studies these factors one at a time, so we know little about how readers weigh them against each other. We address this gap with a choice-based conjoint study (N = 65) in which participants viewed pairs of infographics on a mock newspaper page about unemployment. Each infographic varied across three attributes: comparison type (none, US average, percentage scale), color (red, blue), and graphic type (single icon, icon series, bar chart). Comparison type drove most of the preference variation (58.5%), followed by graphic type (29.2%) and color (12.3%). Readers favored percentage scale markers and benchmark comparisons; color had no practical effect. The percentage scale level adds axis information rather than a benchmark, so the comparison type result mixes two distinct ideas. A single topic and a narrow palette also limit external validity. We argue that conjoint analysis is a practical and underused tool for studying visualization preferences across many design dimensions.

cs.HC

Towards Explainability of SLMs by investigating Token Level Activation

Transformer-based language models such as BERT having 110M+ parameters have revolutionized natural language understanding, yet their internal mechanisms remain largely opaque to researchers and practitioners. Traditional attention-based interpretability methods often emphasize structurally important but semantically weak tokens such as punctuation marks rather than meaningful semantic relationships. This work introduces a lightweight and model-agnostic framework for quantifying token-level representational importance using hidden-state activation strengths at Layer 8 of BERT. The proposed Activation Flow Network (AFN) framework computes Token Activation Strength using the L2 norm of Layer-8 hidden representations, enabling direct ranking of semantically salient tokens. The study further introduces a threshold-based activation bucket formulation that partitions tokens into HIGH-activation and LOW-activation groups using an empirical upper-quartile activation boundary. Experimental observations demonstrate that semantically meaningful content words consistently occupy the HIGH-activation bucket and dominate representational activation shifts, while structurally supportive tokens contribute comparatively less. The results suggest that Layer 8 acts as a critical semantic consolidation zone balancing structural and semantic information processing. By revealing how activation magnitudes concentrate around semantically informative tokens, this work provides an interpretable and computationally efficient alternative to attentioncentric analysis, contributing toward transforming BERT from a "black box" into a more transparent "glass box" model for natural language understanding.

cs.LG

A New Technique for AI Explainability using Feature Association Map

Lack of transparency in AI systems poses challenges in critical real-life applications. It is important to be able to explain the decisions of an AI system to ensure trust on the system. Explainable AI (XAI) algorithms play a vital role in achieving this objective. In this paper, we are proposing a new algorithm for Explaining AI systems, FAMeX (Feature Association Map based eXplainability). The proposed algorithm is based on a graph-theoretic formulation of the feature set termed as Feature Association Map (FAM). The foundation of the modelling is based on association between features. The proposed FAMeX algorithm has been found to be better than the competing XAI algorithms - Permutation Feature Importance (PFI) and SHapley Additive exPlanations (SHAP). Experiments conducted with eight benchmark algorithms show that FAMeX is able to gauge feature importance in the context of classification better than the competing algorithms. This definitely shows that FAMeX is a promising algorithm in explaining the predictions from an AI system

cs.LG

MisVisFix: An Interactive Dashboard for Detecting, Explaining, and Correcting Misleading Visualizations using Large Language Models

Misleading visualizations pose a significant challenge to accurate data interpretation. While recent research has explored the use of Large Language Models (LLMs) for detecting such misinformation, practical tools that also support explanation and correction remain limited. We present MisVisFix, an interactive dashboard that leverages both Claude and GPT models to support the full workflow of detecting, explaining, and correcting misleading visualizations. MisVisFix correctly identifies 96% of visualization issues and addresses all 74 known visualization misinformation types, classifying them as major, minor, or potential concerns. It provides detailed explanations, actionable suggestions, and automatically generates corrected charts. An interactive chat interface allows users to ask about specific chart elements or request modifications. The dashboard adapts to newly emerging misinformation strategies through targeted user interactions. User studies with visualization experts and developers of fact-checking tools show that MisVisFix accurately identifies issues and offers useful suggestions for improvement. By transforming LLM-based detection into an accessible, interactive platform, MisVisFix advances visualization literacy and supports more trustworthy data communication.

cs.HC

Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data Extraction

This paper evaluates the visualization literacy of modern Large Language Models (LLMs) and introduces a novel prompting technique called Charts-of-Thought. We tested three state-of-the-art LLMs (Claude-3.7-sonnet, GPT-4.5 preview, and Gemini-2.0-pro) on the Visualization Literacy Assessment Test (VLAT) using standard prompts and our structured approach. The Charts-of-Thought method guides LLMs through a systematic data extraction, verification, and analysis process before answering visualization questions. Our results show Claude-3.7-sonnet achieved a score of 50.17 using this method, far exceeding the human baseline of 28.82. This approach improved performance across all models, with score increases of 21.8% for GPT-4.5, 9.4% for Gemini-2.0, and 13.5% for Claude-3.7 compared to standard prompting. The performance gains were consistent across original and modified VLAT charts, with Claude correctly answering 100% of questions for several chart types that previously challenged LLMs. Our study reveals that modern multimodal LLMs can surpass human performance on visualization literacy tasks when given the proper analytical framework. These findings establish a new benchmark for LLM visualization literacy and demonstrate the importance of structured prompting strategies for complex visual interpretation tasks. Beyond improving LLM visualization literacy, Charts-of-Thought could also enhance the accessibility of visualizations, potentially benefiting individuals with visual impairments or lower visualization literacy.

cs.HC

Competitive binding of Activator-Repressor in Stochastic Gene Expression

Regulation of gene expression is the consequence of interactions between the promoter of the gene and the transcription factors (TFs). In this paper, we explore the features of a genetic network where the TFs (activators and repressors) bind the promoter in a competitive way. We develop an analytical theory that offers detailed reaction kinetics of the competitive activator-repressor system which could be the powerful tools for extensive study and analysis of the genetic circuit in future research. Moreover, the theoretical approach helps us to find a most probable set of parameter values which was unavailable in experiments. We study the noisy behaviour of the circuit and compare the profile with the network where the activator and repressor bind the promoter non-competitively. We further notice that, due to the effect of transcriptional reinitiation in the presence of the activator and repressor molecules, there exits some anomalous characteristic features in the mean expressions and noise profiles. We find that, in presence of the reinitiation the noise in transcriptional level remains low while it is higher in translational level than the noise when the reinitiation is absent. In addition, it is possible to reduce the noise further below the Poissonian level in competitive circuit than the non-competitive one with the help of some noise reducing parameters.

q-bio.MN

Stochastic gene transcription with non-competitive transcription regulatory architecture

The transcription factors, such as activators and repressors, can interact with the promoter of gene either in a competitive or non-competitive way. In this paper, we construct a stochastic model with non-competitive transcriptional regulatory architecture and develop an analytical theory that re-establishes the experimental results with an improved data fitting. The analytical expressions in the theory allow us to study the nature of the system corresponding to any of its parameters, and hence enable us to find out the factors that govern the regulation of gene expression for that architecture. We notice that, along with transcriptional reinitiation and repressors, there are other parameters that can control the noisiness of this network. We also observe that, the Fano factor (at mRNA level) varies from sub-Poissonian regime to superPoissonian regime. In addition to the aforementioned properties, we observe some anomalous characteristics of the Fano factor (at mRNA level) and that of the variance of protein at lower activator concentrations in presence of repressor molecules. This model is useful to understand the architecture of interactions which may buffer the stochasticity inherent to gene transcription.

q-bio.MN

Bangla hate speech detection on social media using attention-based recurrent neural network

Hate speech has spread more rapidly through the daily use of technology and, most notably, by sharing your opinions or feelings on social media in a negative aspect. Although numerous works have been carried out in detecting hate speeches in English, German, and other languages, very few works have been carried out in the context of the Bengali language. In contrast, millions of people communicate on social media in Bengali. The few existing works that have been carried out need improvements in both accuracy and interpretability. This article proposed encoder decoder based machine learning model, a popular tool in NLP, to classify user's Bengali comments on Facebook pages. A dataset of 7,425 Bengali comments, consisting of seven distinct categories of hate speeches, was used to train and evaluate our model. For extracting and encoding local features from the comments, 1D convolutional layers were used. Finally, the attention mechanism, LSTM, and GRU based decoders have been used for predicting hate speech categories. Among the three encoder decoder algorithms, the attention-based decoder obtained the best accuracy (77%).

cs.CL

Scrutinizing uncitedness of selective Indian physics and astronomy journals through the prism of some h-type indicators

There exist huge chunk of academic items receiving no citation years after years and remaining beyond the veil of ignorance of the academic audience. These are known as uncited items. Now, the question is, why a paper fails to get citation? The attribute of incapability of receiving citation may be termed as Uncitedness. This paper traces brief history of the concept of uncitedness sprouted first in 1964 in an article entitled Cybernetics, homeostasis and a model of disease by Gerson Jacobs. The concept of uncitedness was scientometrically first explained by Garfield in 1970. The uncitedness of twelve esteemed Indian physics and astronomy journals over a twelve years' (2009-2020) time span is analysed here. Besides Uncitedness Factor (UF), three other indicators are introduced here, viz. Citation per paper per Year (CY), h-core Density (HD) and Time-normalised h-index (TH). The journal-wise variational patterns of these four indicators, i.e. UF, CY, HD and TH and the relationships of UF with other three indicators are analysed. The calculated numerical values of these indicators are observed to formulate seven hypotheses, which are tested by F-Test method. The average annual rate of change of uncited paper is found 67% of total number of papers. The indicator CY is found temporally constant. The indicator HD is found nearly constant journal-wise over the entire time span, while the indicator TH is found nearly constant for all journals. The UF inversely varies with CY and TH for the journals and directly varies with TH over the years. Except few highly reputed Indian journals in physics and astronomy, majority other journals face the situation of uncitedness. The uncitedness of Indian journals in this field outshines the same for global journals by 12%, which indicates lack of circulation and timely reach of research communication to the relevant audience.

cs.DL

Effect of transcription reinitiation in stochastic gene expression

Gene expression (GE) is an inherently random or stochastic or noisy process. The randomness in different steps of GE, e.g., transcription, translation, degradation, etc., leading to cell-to-cell variations in mRNA and protein levels. This variation appears in organisms ranging from microbes to metazoans. Stochastic gene expression has important consequences for cellular function. The random fluctuations in protein levels produce variability in cellular behavior. It is beneficial in some contexts and harmful to others. These situations include stress response, metabolism, development, cell cycle, circadian rhythms, and aging. Different model studies e.g., constitutive, two-state, etc., reveal that the fluctuations in mRNA and protein levels arise from different steps of gene expression among which the steps in transcription have the maximum effect. The pulsatile mRNA production through RNAP-II based reinitiation of transcription is an important part of gene transcription. Though, the effect of that process on mRNA and protein levels is very little known. The addition of any biochemical step in the constitutive or two-state process generally decreases the mean and increases the Fano factor. In this study, we have shown that the RNAP-II based reinitiation process in gene transcription can have different effects on both mean and Fano factor at mRNA levels in different model systems. It decreases the mean and Fano factor both at the mRNA levels in the constitutive network whereas in other networks it can simultaneously increase or decrease both quantities or it can have mixed-effect at mRNA levels. We propose that a constitutive network with reinitiation behaves like a product independent negative feedback circuit whereas other networks behave as either product independent positive or negative or mixed feedback circuit.

q-bio.MN