SearcharxivSearch

arXiv subjects

Ayan Chakraborty

Publications and source records attributed to Ayan Chakraborty.

At least 19 recordsLinked to original sources

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference

4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers. Prior work addresses this problem via data rotation or mixed-precision integer quantization, but often relies on software-managed scaling and frequent dequantization, incurring substantial overhead. Microscaling formats, such as MXINT, eliminate these inefficiencies by encoding scales in hardware, yet remain incompatible with rotation-based methods. Our analysis reveals that outliers vary in severity, from rare extremes to frequent mild deviations, and that quantization sensitivity is unevenly distributed across layers and columns. These insights motivate a fine-grained, sensitivity-guided approach. We introduce MXSens, a training-free method that assigns mixed mantissa bitwidths (4/6/8) based on column- and layer-wise sensitivity, naturally leveraging the block-wise structure of MXINT. MXSens outperforms state-of-the-art quantization methods across a range of models and tasks. Under the W4A4KV4 setting, MXSens achieves perplexities of 3.77 and 7.63 on LLaMA-2-70B and LLaMA-3-8B, respectively, substantially improving over existing baselines on WikiText-2. Our work establishes a new balance between accuracy and resource efficiency for LLM quantization.

cs.LG

Constraining Reheating Temperature, Inflaton-SM Coupling and Dark Matter Mass in Light of ACT DR6 Observations

We explore the phenomenological implications of the latest Atacama Cosmology Telescope (ACT) DR6 observations, in combination with Planck 2018, BICEP/Keck 2018, and DESI, on the physics of inflation and post-inflationary reheating. We focus on the $α$-attractor class of inflationary models (both E- and T-models) and consider two reheating scenarios: perturbative inflaton ($ϕ$) decay ($ϕ\rightarrow bb$) and inflaton annihilation ($ϕϕ\rightarrow bb$) into Standard Model (SM) bosonic particles ($b$). By solving the Boltzmann equations, we derive bounds on key reheating parameters, including the reheating temperature, the inflaton equation of state (EoS), and the inflaton-SM coupling, in light of ACT data. To accurately constrain the coupling, we incorporate the Bose enhancement effect in the decay width. To ensure the validity of our perturbative approach, we also identify the regime where nonperturbative effects, such as parametric resonance, become significant. Additionally, we include indirect constraints from primordial gravitational waves (PGWs), which can impact the effective number of relativistic species, $ΔN_{\rm eff}$. These constraints further bound the reheating temperature, particularly in scenarios with a stiff EoS. Finally, we analyze dark matter (DM) production through purely gravitational interactions during reheating and determine the allowed mass ranges consistent with the constrained reheating parameter space and recent ACT data.

hep-ph

Dark matters are Inert, or FIMPy, or WIMPy or UFOy: An inflationary gravitational particle production

In this letter, we explore the phenomenological impact of inflationary gravitational particle production in the physics of Dark Matter (DM). Large-scale DM fluctuations generated during inflation behave as gravitational particles upon their post-inflationary horizon reentry and alter the conventional Boltzmann dynamics of DM with a non-conserving source term, thereby producing significant phenomenological consequences. Within this framework, we analyze four distinct types of DM classified according to their production mechanisms. Dark matter may be completely non-interacting with the thermal bath, behaving as Inert Dark Matter. Alternatively, depending on the strength of its interactions with bath particles, DM may exhibit WIMPy, UFOy, or FIMPy behavior, sharing characteristics with their conventional counterparts. The late-time enhancement of the DM number density, driven by the successive horizon reentry of gravitationally produced low-momentum modes, enlarges the viable parameter space for both thermal and non-thermal DM scenarios. Remarkably, this expanded parameter space remains consistent with current constraints from $ΔN_{\rm eff}$ and Lyman-$α$ bound.

hep-ph

Nonminimal infrared gravitational reheating in light of ACT observation

Inflation is known to produce large infrared scalar fluctuations. Further, if a scalar field $(χ)$ is non-minimally coupled with gravity through $ξχ^2 R$, those infrared modes experience \textit{tachyonic instability} during and after inflation. Those large non-perturbative infrared modes can collectively produce hot Big Bang universe upon their horizon entry during the post-inflationary period. We indeed find that for reheating equation of state (EoS), $w_ϕ > 1/3$, and coupling strength, $ξ>1/6$, large infrared fluctuations lead to successful reheating. We further analyze perturbative reheating by solving the standard Boltzmann equation in both Jordan and Einstein frames, and compare the results with the non-perturbative ones. Finally, embedding this infrared reheating scenario into the well-known $α-$attractor inflationary model, we examine possible constraints on the model parameters in light of the latest ACT, DESI results. To arrive at the constraints, we take into account the latest bounds on tensor-to-scalar ratio, $r_{0.05}\leq 0.038$, isocurvature power spectrum, $\mathcal{P}_{\mathcal{S}} \lesssim 8.3\times 10^{-11}$, and effective number of relativistic degrees of freedom, $ΔN_{\rm eff} \lesssim 0.17 $. Subject to these constraints, we find successful reheating to occur only for EoS $w_ϕ\gtrsim 0.6$, which translates to a sub-class of $α-$attractor models being favored and placing them within the 2$σ$ region in the $ n_s-r$ plane of the latest ACT, DESI data. In this range of EoS, we find that the coupling strength should lie within $2.11\lesssimξ\lesssim 2.95$ for $w_ϕ=0.6$. Finally, we compute secondary gravitational wave signals induced by the scalar infrared modes, which are found to be strong enough to be detected by future GW observatories, namely BBO, DECIGO, LISA, and ET.

astro-ph.CO

Generalizing the Bogoliubov vs Boltzmann approaches in gravitational production

We investigate the spectral behavior of scalar fluctuations generated by gravity during inflation and the subsequent reheating phase. We consider a non-perturbative Bogoliubov treatment within the context of pure gravitational reheating. We compute both long and short-wavelength spectra, first for a massless scalar field, revealing that the spectral index in part of the infrared (IR) regime varies between $-6$ and $-3$, depending on the post-inflationary equation of state (EoS), $0\leq w_ϕ\leq1$. Furthermore, we study the mass-breaking effect of the IR spectrum by including finite mass, $m_χ$, of the daughter scalar field. We show that for $m_χ/H_{\rm e} \gtrsim 3/2$, where $H_{\rm e}$ is the Hubble parameter during inflation, the IR spectrum of scalar fluctuations experiences exponential mass suppression, while for smaller masses, $m_χ/H_{\rm e}<3/2$, the spectrum remains flat in the IR regime regardless of the post-inflationary EoS. For any general EoS, we also compute a specific IR scale, $k_m$, of fluctuations below which the IR spectrum will suffer from this finite mass effect. In the UV regime, oscillations of the inflaton background lead to interference terms that explain the high-frequency oscillations in the spectrum. Interestingly, we find that for any EoS, $1/9 \lesssim w_ϕ\lesssim 1$, the spectral behavior turns out to be independent of the EoS, with a spectral index $ -6$. We have compared this Bogoliubov treatment for the UV regime to perturbative computations with solutions to the Boltzmann equation and found an agreement between the two approaches for any EoS, $0 \lesssim w_ϕ\lesssim 1$. We also explore the relationship between the gravitational reheating temperature and the reheating EoS employing the non-perturbative analytic approach, finding that reheating can occur for $w_ϕ\gtrsim 0.6$.

gr-qc

Probing a nonminimal coupling through superhorizon instability and secondary gravitational waves

In this paper, we investigate the impact of scalar fluctuations ($χ$) non-minimally coupled to gravity, $ξχ^2 R$, as a potential source of secondary gravitational waves (SGWs). Our study reveals that when reheating EoS $\wre < 1/3$ and $ξ\lesssim 1/6$ or $\wre > 1/3$ and $ξ\gtrsim 1/6$, the super-horizon modes of scalar field experience a \textit{Tachyonic instability} during the reheating phase. Such instability causes a substantial growth in the scalar field amplitude leading to pronounced production of SGWs in the low and intermediate-frequency ranges that are strong enough to be detected by Planck and future gravitational wave detectors. Such growth in super-horizon modes of the scalar field and associated GW production may have a significant effect on the strength of the tensor fluctuation at the Cosmic Microwave Background (CMB) scales (parametrized by $r$) and the number of relativistic degrees of freedom (parametrized by $\dneff$) at the time of CMB decoupling. To prevent such overproduction, the PLANCK constraints on tensor-to-scalar ratio $r \leq 0.036$ and $\dneff \leq 0.284$ yield a strong lower bound on $ξ$ for $\wre < 1/3$, and upper bound on the value of $ξ$ for $\wre > 1/3$. Taking into account all the observational constraints we found the value of $ξ$ should be $ \gtrsim 0.02$ for $\wre =0$, and $\lesssim 4.0$ for $\wre \geq 1/2$ for a wide range of reheating temperature within $10^{-2} \lesssim \Tre \lesssim 10^{14}$ GeV, and for a wide range of inflationary energy scales. Further, as one approaches $\wre$ towards $1/3$, the value of $ξ$ remains unconstrained. Finally, we identify the parameter regions in $(\Tre,ξ)$ plane which can be probed by the upcoming GW experiments namely BBO, DECIGO, LISA, and ET.

astro-ph.CO

Effective Interplay between Sparsity and Quantization: From Theory to Practice

The increasing size of deep neural networks (DNNs) necessitates effective model compression to reduce their computational and memory footprints. Sparsity and quantization are two prominent compression methods that have been shown to reduce DNNs' computational and memory footprints significantly while preserving model accuracy. However, how these two methods interact when combined together remains a key question for developers, as many tacitly assume that they are orthogonal, meaning that their combined use does not introduce additional errors beyond those introduced by each method independently. In this paper, we provide the first mathematical proof that sparsity and quantization are non-orthogonal. We corroborate these results with experiments spanning a range of large language models, including the OPT and LLaMA model families (with 125M to 8B parameters), and vision models like ViT and ResNet. We show that the order in which we apply these methods matters because applying quantization before sparsity may disrupt the relative importance of tensor elements, which may inadvertently remove significant elements from a tensor. More importantly, we show that even if applied in the correct order, the compounded errors from sparsity and quantization can significantly harm accuracy. Our findings extend to the efficient deployment of large models in resource-constrained compute platforms to reduce serving cost, offering insights into best practices for applying these compression methods to maximize hardware resource efficiency without compromising accuracy.

cs.LG

Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training

The unprecedented demand for computing resources to train DNN models has led to a search for minimal numerical encoding. Recent state-of-the-art (SOTA) proposals advocate for multi-level scaled narrow bitwidth numerical formats. In this paper, we show that single-level scaling is sufficient to maintain training accuracy while maximizing arithmetic density. We identify a previously proposed single-level scaled format for 8-bit training, Hybrid Block Floating Point (HBFP), as the optimal candidate to minimize. We perform a full-scale exploration of the HBFP design space using mathematical tools to study the interplay among various parameters and identify opportunities for even smaller encodings across layers and epochs. Based on our findings, we propose Accuracy Booster, a mixed-mantissa HBFP technique that uses 4-bit mantissas for over 99% of all arithmetic operations in training and 6-bit mantissas only in the last epoch and first/last layers. We show Accuracy Booster enables increasing arithmetic density over all other SOTA formats by at least 2.3x while achieving state-of-the-art accuracies in 4-bit training.

cs.LG

A Kalman Filter based Low Complexity Throughput Prediction Algorithm for 5G Cellular Networks

Throughput Prediction is one of the primary preconditions for the uninterrupted operation of several network-aware mobile applications, namely video streaming. Recent works have advocated using Machine Learning (ML) and Deep Learning (DL) for cellular network throughput prediction. In contrast, this work has proposed a low computationally complex simple solution which models the future throughput as a multiple linear regression of several present network parameters and present throughput. It then feeds the variance of prediction error and measurement error, which is inherent in any measurement setup but unaccounted for in existing works, to a Kalman filter-based prediction-correction approach to obtain the optimal estimates of the future throughput. Extensive experiments across seven publicly available 5G throughput datasets for different prediction window lengths have shown that the proposed method outperforms the baseline ML and DL algorithms by delivering more accurate results within a shorter timeframe for inferencing and retraining. Furthermore, in comparison to its ML and DL counterparts, the proposed throughput prediction method is also found to deliver higher QoE to both streaming and live video users when used in conjunction with popular Model Predictive Control (MPC) based adaptive bitrate streaming algorithms.

cs.NI

Squeezing, Chaos and Thermalization in Periodically Driven Quantum Systems: The Case of Bosonic Preheating

The phenomena of Squeezing and chaos have recently been studied in the context of inflation. We apply this formalism in the post-inflationary preheating phase. During this phase, inflaton field undergoes quasi-periodic oscillation, which acts as a driving force for the resonant growth of quantum fluctuation or particle production. Furthermore, the quantum state of the fluctuations is known to have evolved into a squeezed state. In this submission, we explore the underlying connection between the resonant growth, squeezing, and chaos by computing the Out of Time Order Correlator (OTOC) of phase space variables and establishing a relation among the Lyapunov, Floquet exponents, and squeezing parameters. For our study, we consider observationally favored $α$-attractor E-model of inflaton which is coupled with the bosonic field. After the production, the system of produced bosonic fluctuations/particles from the inflaton is supposed to thermalize, and that is believed to have an intriguing connection to the nature of chaos of the system under perturbation. %By using this we calculated approximate lower bound of temperature ${\bar T}_{\rm MSS}$. We conjecture a relation between the thermalization temperature $({\bar T}_{\rm SS})$ of the system and quantum squeezing, which is further shown to be consistent with the well-known Rayleigh-Jeans formula for the temperature symbolized as ${\bar T}_{\rm RJ}$, and that is ${\bar T}_{\rm SS} \simeq {\bar T}_{\rm RJ}$. Finally, we show that the system temperature is in accord with the well-known lower bound on the temperature of a chaotic system proposed by Maldacena-Shenker-Stanford (MSS).

hep-th

Inflaton phenomenology via reheating in light of primordial gravitational waves and the latest BICEP/$Keck$ data

We are in the era of precision cosmology which offers us a unique opportunity to investigate beyond standard model physics. Toward this endeavor, inflaton is assumed to be a perfect new physics candidate. In this submission, we explore the phenomenological impact of the latest observation of PLANCK and BICEP/$Keck$ data on the physics of inflation. We particularly study three different models of inflation, namely $α$-attractor E, T, and the minimal plateau model. We further consider two different post-inflationary reheating dynamics driven by inflaton decaying into bosons and fermions. Given the latest data in the inflationary $(n_s-r)$ plane, we derive detailed phenomenological constraints on different inflaton parameters and the associated physical quantities, such as inflationary $e$-folding number, $N_{ k}$, reheating temperatures $T_{\rm re}$. Apart from considering direct observational data, we further incorporate the bounds from primordial gravitational waves (PGWs) and different theoretical constraints. Rather than in the laboratory, our results illustrate the potential of present and future cosmological observations to look for new physics in the sky.

astro-ph.CO

Variational energy based XPINNs for phase field analysis in brittle fracture

Modeling fracture is computationally expensive even in computational simulations of two-dimensional problems. Hence, scaling up the available approaches to be directly applied to large components or systems crucial for real applications become challenging. In this work. we propose domain decomposition framework for the variational physics-informed neural networks to accurately approximate the crack path defined using the phase field approach. We show that coupling domain decomposition and adaptive refinement schemes permits to focus the numerical effort where it is most needed: around the zones where crack propagates. No a priori knowledge of the damage pattern is required. The ability to use numerous deep or shallow neural networks in the smaller subdomains gives the proposed method the ability to be parallelized. Additionally, the framework is integrated with adaptive non-linear activation functions which enhance the learning ability of the networks, and results in faster convergence. The efficiency of the proposed approach is demonstrated numerically with three examples relevant to engineering fracture mechanics. Upon the acceptance of the manuscript, all the codes associated with the manuscript will be made available on Github.

cs.CE

Multigoal-oriented dual-weighted-residual error estimation using deep neural networks

Deep learning has shown successful application in visual recognition and certain artificial intelligence tasks. Deep learning is also considered as a powerful tool with high flexibility to approximate functions. In the present work, functions with desired properties are devised to approximate the solutions of PDEs. Our approach is based on a posteriori error estimation in which the adjoint problem is solved for the error localization to formulate an error estimator within the framework of neural network. An efficient and easy to implement algorithm is developed to obtain a posteriori error estimate for multiple goal functionals by employing the dual-weighted residual approach, which is followed by the computation of both primal and adjoint solutions using the neural network. The present study shows that such a data-driven model based learning has superior approximation of quantities of interest even with relatively less training data. The novel algorithmic developments are substantiated with numerical test examples. The advantages of using deep neural network over the shallow neural network are demonstrated and the convergence enhancing techniques are also presented

cs.LG

Methodology for Biasing Random Simulation for Rapid Coverage of Corner Cases in AMS Designs

Exploring the limits of an Analog and Mixed Signal (AMS) circuit by driving appropriate inputs has been a serious challenge to the industry. Doing an exhaustive search of the entire input state space is a time-consuming exercise and the returns to efforts ratio is quite low. In order to meet time-to-market requirements, often suboptimal coverage results of an integrated circuit (IC) are leveraged. Additionally, no standards have been defined which can be used to identify a target in the continuous state space of analog domain such that the searching algorithm can be guided with some heuristics. In this report, we elaborate on two approaches for tackling this challenge - one is based on frequency domain analysis of the circuit, while the other applies the concept of Bayesian optimization. We have also presented our results by applying the two approaches on an industrial LDO and a few AMS benchmark circuits.

cs.FL

Synthesis of Feedback Controller for Nonlinear Control Systems with Optimal Region of Attraction

We propose a framework for synthesizing a feedback control policy that maximizes the region of attraction (ROA) of a closed-loop nonlinear dynamical system. Our synthesis technique relies on stochastic optimization, which involves computation of an objective function capturing the ROA for a feedback control law. We employ a machine learning technique based on deep neural network to estimate the ROA for a given feedback controller. Overall, our technique is capable of synthesizing a controller co-optimizing traditional control objectives like LQR cost together with ROA. We demonstrate the efficacy of our technique through exhaustive experiments carried out on various nonlinear systems.

eess.SY

Non uniform weighted extended B-Spline finite element analysis of non linear elliptic partial differential equations

We propose a non uniform web spline based finite element analysis for elliptic partial differential equation with the gradient type nonlinearity in their principal coefficients like p-laplacian equation and Quasi-Newtonian fluid flow equations. We discuss the well-posednes of the problems and also derive the apriori error estimates for the proposed finite element analysis and obtain convergence rate of $\mathcal{O}(h^α)$ for $α> 0$.

math.NA

Weighted Extended B-Spline Finite Element Analysis of a coupled system of general Elliptic equations

In this study we establish the existence and uniqueness of the solution of a coupled system of general elliptic equations with anisotropic diffusion , non-uniform advection and variably influencing reaction terms on Lipschitz continuous domain $Ω\subset \mathbb{R}^m $ (m$\geq$1) with a Dirichlet boundary. Later we consider the finite element (FE) approximation of the coupled equations in a meshless framework based on weighted extended B-Spine functions (WEBS).The a priori error estimates corresponding to the finite element analysis are derived to establish the convergence of the corresponding FE scheme and the numerical methodology has been tested on few examples.

math.NA

Web spline error estimation of non-cooperative elliptic equations for population dynamics

We analyze the error of the WEB-S finite element method applied to elliptic systems with non-cooperative dominant coupling,with a mixed Dirichlet/Neumann/Robin boundary condition. This problem is strongly related to a posteriori error estimates, giving computable bounds for computational errors and detecting zones in the solution domain where such errors are too large and certain mesh refinements should be performed. These results are based on an extensive regularity analysis of the interface problems of concern.Finally, the error analysis is illustrated by numerical experiments.

math.NA