SearcharxivSearch

arXiv subjects

Sumanta Mukherjee

Publications and source records attributed to Sumanta Mukherjee.

At least 19 recordsLinked to original sources

Can a Spin Liquid State Persist as the Ground State in the Presence of Competing Interactions and Disorder?

In this report, we have shown that, in an otherwise geometrically frustrated lattice, the underlying interactions compete to stabilize different magnetic phases. In conjunction with fluctuations, these interactions may lead to the formation of unusual ordered phases, providing a pathway for understanding the fluctuation-driven order-by-disorder phenomenon. This competition is further modified by the presence of inherent structural disorder in a frustrated two-dimensional lattice. Furthermore, due to the local nature of the additional structural disorder, the majority of the samples evolve toward glassy dynamics, which can not only mimic spin-liquid-like behavior but also make it difficult to identify genuine spin-liquid candidates experimentally.

cond-mat.mtrl-sci

Quantum Variational Approaches to the Maximum Independent Set Problem at Utility Scale

We study variational quantum algorithms for the Maximum Independent Set (MIS) problem on benchmark graphs of 64, 99, and 180 vertices. The Variational Quantum Eigensolver (VQE) and Quantum Approximate Optimization Algorithm (QAOA) are compared across SPSA and COBYLA optimizers at multiple circuit depths. A preprocessing pipeline comprising spectral graph reordering (via the Fiedler vector) and distance-based sparsification reduces circuit depth while preserving energy fidelity. Classical post-processing via history-guided bitstring correction and stepwise maximality extension recovers the exact MIS across all instances. With CVaR optimization, VQE with SPSArecovers up to 6 distinct MIS per run for the 64-node instance and up to 10 distinct MIS per run for the 99-node instance, sampling broadly from the optimal solution population. Repeated runs with different SPSA trajectories collectively enumerate a larger fraction of all MIS for each instance. For the 180-node instance, where standard approaches stall at size 14 (MIS is 15), we introduce ancilla-assisted superposition initialization: ancilla qubits prepare a uniform superposition over classically-found near-optimal solutions, and an excitation-preserving ansatz evolves this state while conserving Hamming weight. This novel construction enables quantum-parallel variational search over multiple seeds simultaneously, discovering the exact MIS where single-seed methods fail. The 180-qubit simulation represents, to our knowledge, the largest scale at which gate-based variational algorithms have solved MIS to optimality. Hardware validation on IBM Quantum hardware ibm_marrakesh confirms that converged simulator parameters transfer effectively to noisy quantum execution.

quant-ph

Hamiltonian-Guided Leverage Embedding: Robust Subspace Compression for Efficient QAOA Parameter Estimation

The Quantum Approximate Optimization Algorithm (QAOA) is a hybrid quantum-classical framework for combinatorial optimization on near-term quantum devices. A central bottleneck is the classical estimation of its variational parameters {\gamma} and {\beta}, which must be optimized over a high-dimensional, non-convex landscape corrupted by sampling noise. We observe that the classical feature matrices constructed from QAOA measurement samples exhibit pronounced low-rank structure, and exploit this property for noise-robust, reduced-dimension parameter search. We present the Hamiltonian-Guided Leverage Embedding (HGLE) algorithm - a hybrid pipeline that encodes low-energy quantum samples into a weighted Ising feature matrix and compresses it via leverage-score row sampling, provably preserving the dominant rank-rsubspace geometry. The compressed representation drives a classical trust-region loop for ({\gamma}, {\beta}) estimation at a fraction of the original cost. We provide formal guarantees for rank preservation and energy approximation error, and demonstrate robustness across problem types (Max-Cut, Maximum Independent Set) and graph topologies of varying density.

quant-ph

A unified equation for saturation magnetization and spin transport in weakly disordered ferromagnets

In this report, a unified description of the loss of saturation magnetization in the presence of a distribution of finite-size effects is provided for weakly disordered spin-1/2 ferromagnets. This description allows us to derive a unified form of the Bloch equation for these systems. We further extend this approach to obtain a unified expression for spin transport in such systems.

cond-mat.mtrl-sci

Magnetic Behavior of Ferro-, Antiferro-, and Ferrimagnetic Systems in the Griffiths Phase: A Theoretical Study

In this report, we provide a theoretical framework for the magnetic behavior of the Griffiths phase, which, along with three-dimensional spin-1/2 Ising ferromagnetic systems, can be extended to antiferromagnetic as well as ferrimagnetic systems. We find that the magnetic behavior in the Griffiths phase of three-dimensional antiferromagnetic and ferrimagnetic systems is more unusual than that of conventional ferromagnetic systems. However, this study offers a possible framework for the identification of Griffiths phase behavior in three-dimensional antiferromagnetic and ferrimagnetic systems.

cond-mat.mtrl-sci

Spontaneous Emission, Work Potential and Relaxation-Limited Processes in Setting Limits on Solar Energy Conversion Efficiency

Understanding the thermodynamics of radiation and the quantum-mechanical interactions between light and matter is important both for theoretical purposes and for technological advances, such as determining the limits of key processes like light-to-usable-energy conversion efficiencies. In this report, we discuss the physics of these two aspects, considering spontaneous emission as a pathway, and highlight the limitations of such descriptions in assessing energy-harvesting efficiency. In view of these limitations, we adopt a simplified approach to evaluate the exergy and work potential of radiation, providing a framework for assessing various aspects of light-to-usable-energy conversion efficiency. Our approach allows a theoretical estimate of the thermodynamic maximum limit for light-to-usable-energy conversion, which is approximately 76%. We validate these exergy and work potential estimates by modeling and accurately reproducing the Shockley-Queisser limit (~ 33.3%), which imposes a practical constraint on solar-to-usable-energy conversion efficiency. Beyond exergy considerations, our model incorporates processes such as spontaneous emission, nonradiative thermal losses, and photon upconversion, allowing us to evaluate their roles. The model further suggests that, under certain conditions, the maximum conversion efficiency can reach approximately 48%, for example with multijunction solar cells or via photon upconversion. These findings further suggest that the true thermodynamic limit for light-to-usable-energy conversion may be much higher (approximately 76%). However, accurately estimating this limit requires a more complete understanding of the thermodynamics of light, light-matter interactions, and the connection between them.

physics.app-ph

TSPulse: Tiny Pre-Trained Models with Disentangled Representations for Rapid Time-Series Analysis

Time-series tasks often benefit from signals expressed across multiple representation spaces (e.g., time vs. frequency) and at varying abstraction levels (e.g., local patterns vs. global semantics). However, existing pre-trained time-series models entangle these heterogeneous signals into a single large embedding, limiting transferability and direct zero-shot usability. To address this, we propose TSPulse, family of ultra-light pre-trained models (1M parameters) with disentanglement properties, specialized for various time-series diagnostic tasks. TSPulse introduces a novel pre-training framework that augments masked reconstruction with explicit disentanglement across spaces and abstractions, learning three complementary embedding views (temporal, spectral, and semantic) to effectively enable zero-shot transfer. In-addition, we introduce various lightweight post-hoc fusers that selectively attend and fuse these disentangled views based on task type, enabling simple but effective task specializations. To further improve robustness and mitigate mask-induced bias prevalent in existing approaches, we propose a simple yet effective hybrid masking strategy that enhances missing diversity during pre-training. Despite its compact size, TSPulse achieves strong and consistent gains across four TS diagnostic tasks: +20% on the TSB-AD anomaly detection leaderboard, +25% on similarity search, +50% on imputation, and +5-16% on multivariate classification, outperforming models that are 10-100X larger on over 75 datasets. TSPulse delivers state-of-the-art zero-shot performance, efficient fine-tuning, and supports GPU-free deployment. Models and source code are publicly available at https://huggingface.co/ibm-granite/granite-timeseries-tspulse-r1.

cs.LG

Percolation in a three-dimensional non-symmetric multi-color loop model

We conducted Monte Carlo simulations to analyze the percolation transition of a non-symmetric loop model on a regular three-dimensional lattice. We calculated the critical exponents for the percolation transition of this model. The percolation transition occurs at a temperature that is close to, but not exactly the thermal critical temperature. Our finite-size study on this model yielded a correlation length exponent that agrees with that of the three-dimensional XY model with an error margin of six per cent.

cond-mat.stat-mech

Thermodynamics of multi-colored loop models in three dimensions

We study order-disorder transitions in three-dimensional \textsl{multi-colored} loop models using Monte Carlo simulations. We show that the nature of the transition is intimately related to the nature of the loops. The symmetric loops undergo a first order phase transition, while the non-symmetric loops show a second-order transition. The critical exponents for the non-symmetric loops are calculated. In three dimensions, the regular loop model with no interactions is dual to the XY model. We argue that, due to interactions among the colors, the specific heat exponent is found to be different from that of the regular loop model. The continuous nature of the transition is altered to a discontinuous one due to the strong inter-color interactions.

cond-mat.stat-mech

Towards Unbiased Evaluation of Time-series Anomaly Detector

Time series anomaly detection (TSAD) is an evolving area of research motivated by its critical applications, such as detecting seismic activity, sensor failures in industrial plants, predicting crashes in the stock market, and so on. Across domains, anomalies occur significantly less frequently than normal data, making the F1-score the most commonly adopted metric for anomaly detection. However, in the case of time series, it is not straightforward to use standard F1-score because of the dissociation between `time points' and `time events'. To accommodate this, anomaly predictions are adjusted, called as point adjustment (PA), before the $F_1$-score evaluation. However, these adjustments are heuristics-based, and biased towards true positive detection, resulting in over-estimated detector performance. In this work, we propose an alternative adjustment protocol called ``Balanced point adjustment'' (BA). It addresses the limitations of existing point adjustment methods and provides guarantees of fairness backed by axiomatic definitions of TSAD evaluation.

cs.LG

Activations Through Extensions: A Framework To Boost Performance Of Neural Networks

Activation functions are non-linearities in neural networks that allow them to learn complex mapping between inputs and outputs. Typical choices for activation functions are ReLU, Tanh, Sigmoid etc., where the choice generally depends on the application domain. In this work, we propose a framework/strategy that unifies several works on activation functions and theoretically explains the performance benefits of these works. We also propose novel techniques that originate from the framework and allow us to obtain ``extensions'' (i.e. special generalizations of a given neural network) of neural networks through operations on activation functions. We theoretically and empirically show that ``extensions'' of neural networks have performance benefits compared to vanilla neural networks with insignificant space and time complexity costs on standard test functions. We also show the benefits of neural network ``extensions'' in the time-series domain on real-world datasets.

cs.LG

Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series

Large pre-trained models excel in zero/few-shot learning for language and vision tasks but face challenges in multivariate time series (TS) forecasting due to diverse data characteristics. Consequently, recent research efforts have focused on developing pre-trained TS forecasting models. These models, whether built from scratch or adapted from large language models (LLMs), excel in zero/few-shot forecasting tasks. However, they are limited by slow performance, high computational demands, and neglect of cross-channel and exogenous correlations. To address this, we introduce Tiny Time Mixers (TTM), a compact model (starting from 1M parameters) with effective transfer learning capabilities, trained exclusively on public TS datasets. TTM, based on the light-weight TSMixer architecture, incorporates innovations like adaptive patching, diverse resolution sampling, and resolution prefix tuning to handle pre-training on varied dataset resolutions with minimal model capacity. Additionally, it employs multi-level modeling to capture channel correlations and infuse exogenous signals during fine-tuning. TTM outperforms existing popular benchmarks in zero/few-shot forecasting by (4-40%), while reducing computational requirements significantly. Moreover, TTMs are lightweight and can be executed even on CPU-only machines, enhancing usability and fostering wider adoption in resource-constrained environments. The model weights for reproducibility and research use are available at https://huggingface.co/ibm/ttm-research-r2/, while enterprise-use weights under the Apache license can be accessed as follows: the initial TTM-Q variant at https://huggingface.co/ibm-granite/granite-timeseries-ttm-r1, and the latest variants (TTM-B, TTM-E, TTM-A) weights are available at https://huggingface.co/ibm-granite/granite-timeseries-ttm-r2.

cs.LG

TsSHAP: Robust model agnostic feature-based explainability for time series forecasting

A trustworthy machine learning model should be accurate as well as explainable. Understanding why a model makes a certain decision defines the notion of explainability. While various flavors of explainability have been well-studied in supervised learning paradigms like classification and regression, literature on explainability for time series forecasting is relatively scarce. In this paper, we propose a feature-based explainability algorithm, TsSHAP, that can explain the forecast of any black-box forecasting model. The method is agnostic of the forecasting model and can provide explanations for a forecast in terms of interpretable features defined by the user a prior. The explanations are in terms of the SHAP values obtained by applying the TreeSHAP algorithm on a surrogate model that learns a mapping between the interpretable feature space and the forecast of the black-box model. Moreover, we formalize the notion of local, semi-local, and global explanations in the context of time series forecasting, which can be useful in several scenarios. We validate the efficacy and robustness of TsSHAP through extensive experiments on multiple datasets.

cs.LG

Semi-supervised counterfactual explanations

Counterfactual explanations for machine learning models are used to find minimal interventions to the feature values such that the model changes the prediction to a different output or a target output. A valid counterfactual explanation should have likely feature values. Here, we address the challenge of generating counterfactual explanations that lie in the same data distribution as that of the training data and more importantly, they belong to the target class distribution. This requirement has been addressed through the incorporation of auto-encoder reconstruction loss in the counterfactual search process. Connecting the output behavior of the classifier to the latent space of the auto-encoder has further improved the speed of the counterfactual search process and the interpretability of the resulting counterfactual explanations. Continuing this line of research, we show further improvement in the interpretability of counterfactual explanations when the auto-encoder is trained in a semi-supervised fashion with class tagged input data. We empirically evaluate our approach on several datasets and show considerable improvement in-terms of several metrics.

cs.LG

Hierarchical Proxy Modeling for Improved HPO in Time Series Forecasting

Selecting the right set of hyperparameters is crucial in time series forecasting. The classical temporal cross-validation framework for hyperparameter optimization (HPO) often leads to poor test performance because of a possible mismatch between validation and test periods. To address this test-validation mismatch, we propose a novel technique, H-Pro to drive HPO via test proxies by exploiting data hierarchies often associated with time series datasets. Since higher-level aggregated time series often show less irregularity and better predictability as compared to the lowest-level time series which can be sparse and intermittent, we optimize the hyperparameters of the lowest-level base-forecaster by leveraging the proxy forecasts for the test period generated from the forecasters at higher levels. H-Pro can be applied on any off-the-shelf machine learning model to perform HPO. We validate the efficacy of our technique with extensive empirical evaluation on five publicly available hierarchical forecasting datasets. Our approach outperforms existing state-of-the-art methods in Tourism, Wiki, and Traffic datasets, and achieves competitive result in Tourism-L dataset, without any model-specific enhancements. Moreover, our method outperforms the winning method of the M5 forecast accuracy competition.

cs.LG

Intersection Patterns in Optimal Binary $(5,3)$ Doubling Subspace Codes

Subspace codes are collections of subspaces of a projective space such that any two subspaces satisfy a pairwise minimum distance criterion. Recent results have shown that it is possible to construct optimal $(5,3)$ subspace codes from pairs of partial spreads in the projective space $\mathrm{PG}(4,q)$ over the finite field $ \mathbb{F}_q $, termed doubling codes. We have utilized a complete classification of maximal partial line spreads in $\mathrm{PG}(4,2)$ in literature to establish the types of the spreads in the doubling code instances obtained from two recent constructions of optimum $(5,3)_q$ codes, restricted to $ \mathbb{F}_2 $. Further we present a new characterization of a subclass of binary doubling codes based on the intersection patterns of key subspaces in the pair of constituent spreads.

cs.IT

Explainable AI based Interventions for Pre-season Decision Making in Fashion Retail

Future of sustainable fashion lies in adoption of AI for a better understanding of consumer shopping behaviour and using this understanding to further optimize product design, development and sourcing to finally reduce the probability of overproducing inventory. Explainability and interpretability are highly effective in increasing the adoption of AI based tools in creative domains like fashion. In a fashion house, stakeholders like buyers, merchandisers and financial planners have a more quantitative approach towards decision making with primary goals of high sales and reduced dead inventory. Whereas, designers have a more intuitive approach based on observing market trends, social media and runways shows. Our goal is to build an explainable new product forecasting tool with capabilities of interventional analysis such that all the stakeholders (with competing goals) can participate in collaborative decision making process of new product design, development and launch.

cs.CY

Semi-Supervised Method using Gaussian Random Fields for Boilerplate Removal in Web Browsers

Boilerplate removal refers to the problem of removing noisy content from a webpage such as ads and extracting relevant content that can be used by various services. This can be useful in several features in web browsers such as ad blocking, accessibility tools such as read out loud, translation, summarization etc. In order to create a training dataset to train a model for boilerplate detection and removal, labeling or tagging webpage data manually can be tedious and time consuming. Hence, a semi-supervised model, in which some of the webpage elements are labeled manually and labels for others are inferred based on some parameters, can be useful. In this paper we present a solution for extraction of relevant content from a webpage that relies on semi-supervised learning using Gaussian Random Fields. We first represent the webpage as a graph, with text elements as nodes and the edge weights representing similarity between nodes. After this, we label a few nodes in the graph using heuristics and label the remaining nodes by a weighted measure of similarity to the already labeled nodes. We describe the system architecture and a few preliminary results on a dataset of webpages.

stat.ML