SearcharxivSearch

arXiv subjects

Sourish Das

Publications and source records attributed to Sourish Das.

At least 19 recordsLinked to original sources

TinyBayes: Closed-Form Bayesian Inference via Jacobi Prior for Real-Time Image Classification on Edge Devices

Cocoa (Theobroma cacao) is a critical cash crop for millions of smallholder farmers in West Africa, where Cocoa Swollen Shoot Virus Disease (CSSVD) and anthracnose cause devastating yield losses. Automated disease detection from leaf images is essential for early intervention, yet deploying such systems in resource-constrained settings demands models that are small, fast, and require no internet connectivity. Existing edge-deployable plant disease systems rely on end-to-end deep learning without uncertainty quantification, while Bayesian methods for edge devices focus on hardware-level inference architectures rather than agricultural applications. We bridge this gap with TinyBayes, the first framework to combine a closed-form Bayesian classifier with a mobile-grade computer vision pipeline for crop disease detection. Our pipeline uses YOLOv8-Nano (5.9 MB) for lesion localisation, MobileNetV3-Small (3.5 MB) for feature extraction, and the Jacobi prior; a Bayesian method that provides a closed form non-iterative estimators via projection, for the classification. The Jacobi-DMR (Distributed Multinomial Regression) classifier adds only 13.5 KB to the pipeline, bringing the total model size within 9.5 MB, while achieving 78.7% accuracy on the Amini Cocoa Contamination Challenge dataset and enabling end-to-end CPU inference under 150 ms per image. We benchmark against seven classifiers including Random Forest, SVM, Ridge, Lasso, Elastic Net, XGBoost, and Jacobi-GP, and demonstrate that the Jacobi-DMR offers the best trade-off between accuracy, model size, and inference speed for edge deployment. We have proved the asymptotic equivalence and consistency, asymptotic normality and the bias correction of Jacobi-DMR. All data and codes are available here: https://github.com/shouvik-sardar/TinyBayes

cs.CV

Jacobi Prior: An Alternative Bayesian Method for Supervised Learning

The Jacobi prior offers an alternative Bayesian framework, designed to achieve superior computational efficiency without compromising predictive performance. Compared to widely used methods such as Lasso, Ridge, Elastic Net, uniLasso, the MCMC-based Horseshoe prior, and non-Bayesian machine learning methods including Support Vector Machines (SVM), Random Forests, and Extreme Gradient Boosting (XGBoost), the Jacobi prior achieves competitive or better accuracy with significantly reduced computational cost. The method is well suited to distributed computing environments, as it naturally accommodates partitioned data across multiple servers. We propose a parallelisable Monte Carlo algorithm to quantify the uncertainty in the estimated coefficients. We establish that the Jacobi estimator is asymptotically close to, and asymptotically equivalent to, the posterior mode under the Jacobi prior. To demonstrate its practical utility, we conduct a comprehensive simulation study comprising seven experiments focused on statistical consistency, prediction accuracy, scalability, sensitivity analysis and robustness study. We further present three real-data applications multi-class classification of stars, quasars, and galaxies using Sloan Digital Sky Survey data, and spinal degeneration classification using sagittal MRI scans from the RSNA 2024 Lumbar Spine Degenerative Classification Challenge. In the spine classification task, we extract last-layer features from a fine-tuned ResNet-50 model and evaluate multiple classifiers, including Jacobi-Multinomial logit regression, SVM, and Random Forest. All code and datasets used in this paper are available at: https://github.com/sourish-cmi/Jacobi-Prior/

stat.ME

The Impact of Meteorological Factors on Crop Price Volatility in India: Case studies of Soybean and Brinjal

Climate is an evolving complex system with dynamic interactions and non-linear feedback mechanisms, shaping environmental and socio-economic outcomes. Crop production is highly sensitive to climatic fluctuations (and many other environmental, social and governance factors). This paper studies the price volatility of agricultural crops as influenced by meteorological variables, which is critical for agricultural planning, sustainable finance and policy-making. As case studies, we choose the two Indian states: Madhya Pradesh (for Soybean) and Odisha (for Brinjal/Eggplant). We employ an Exponential Generalized Autoregressive Conditional Heteroskedasticity (EGARCH) model to estimate the conditional volatility of the log returns from 2012 to 2024. We further explore the cross-correlations between price volatility and the meteorological variables followed by a Granger-causal test to analyze the causal effect of meteorological variables on the volatility. The Seasonal Auto-Regressive Integrated Moving Average with Exogenous Regressors (SARIMAX) and Long Short-Term Memory (LSTM) models are implemented as simple machine learning models of price volatility with meteorological factors as exogenous variables. Finally, to capture spatial dependencies in volatility across districts, we extend the analysis using a Conditional Autoregressive (CAR) model to construct monthly volatility surfaces that reflect both local price risk as well as geographic dependence. We believe, this paper will illustrate the usefulness of simple machine learning models in agricultural finance, and help the farmers to make informed decisions by considering climate patterns and making beneficial decisions with regard to crop rotation or allocations. In general, incorporating meteorological factors to assess agricultural performance could help to understand and reduce price volatility and possibly lead to economic stability.

stat.AP

Predicting Stock Market Crash with Bayesian Generalised Pareto Regression

This paper develops a Bayesian Generalised Pareto Regression (GPR) model to forecast extreme losses in Indian equity markets, with a focus on the Nifty 50 index. Extreme negative returns, though rare, can cause significant financial disruption, and accurate modelling of such events is essential for effective risk management. Traditional Generalised Pareto Distribution (GPD) models often ignore market conditions; in contrast, our framework links the scale parameter to covariates using a log-linear function, allowing tail risk to respond dynamically to market volatility. We examine four prior choices for Bayesian regularisation of regression coefficients: Cauchy, Lasso (Laplace), Ridge (Gaussian), and Zellner's g-prior. Simulation results suggest that the Cauchy prior delivers the best trade-off between predictive accuracy and model simplicity, achieving the lowest RMSE, AIC, and BIC values. Empirically, we apply the model to large negative returns (exceeding 5%) in the Nifty 50 index. Volatility measures from the Nifty 50, S&P 500, and gold are used as covariates to capture both domestic and global risk drivers. Our findings show that tail risk increases significantly with higher market volatility. In particular, both S&P 500 and gold volatilities contribute meaningfully to crash prediction, highlighting global spillover and flight-to-safety effects. The proposed GPR model offers a robust and interpretable approach for tail risk forecasting in emerging markets. It improves upon traditional EVT-based models by incorporating real-time financial indicators, making it useful for practitioners, policymakers, and financial regulators concerned with systemic risk and stress testing.

q-fin.ST

Mitigating Financial Risk from Climate-Induced Agricultural Price Volatility

Agricultural price volatility, driven by market dynamics and meteorological factors such as temperature and precipitation, poses challenges for sustainable finance, planning, and policy. This study analyzes the impact of climate on crop price volatility for soybean in Madhya Pradesh (India) and Illinois (US), rice in Assam (India), wheat in North Dakota (US), cotton in Gujarat (India), and corn in Iowa (US). Using CMIP6 climate projections from the Copernicus Climate Change Service, we examine historical climate patterns and evaluate two future scenarios: SSP2-4.5 (moderate) and SSP5-8.5 (severe). We estimate conditional price volatility using the Exponential Generalized Autoregressive Conditional Heteroskedasticity (EGARCH) model, and forecast this volatility with a Seasonal Autoregressive Integrated Moving Average with Exogenous Regressors (SARIMAX) model that incorporates meteorological variables. Finally, we apply the Black-Scholes framework to evaluate the cost of put-option-based insurance, which provides protection to farmers against adverse price drops linked to climate change. Our results highlight the role of meteorological data in improving agricultural risk modelling, enabling better design of insurance mechanisms, price stabilization tools, and sustainable policy interventions under climate uncertainty.

stat.AP

Causal Links Between Anthropogenic Emissions and Air Pollution Dynamics in Delhi

Air pollution poses significant health and environmental challenges, particularly in rapidly urbanizing regions. Delhi-National Capital Region experiences air pollution episodes due to complex interactions between anthropogenic emissions and meteorological conditions. Understanding the causal drivers of key pollutants such as $PM_{2.5}$ and ground $O_3$ is crucial for developing effective mitigation strategies. This study investigates the causal links of anthropogenic emissions on $PM_{2.5}$ and $O_3$ concentrations using predictive modeling and causal inference techniques. Integrating high-resolution air quality data from Jan 2018 to Aug 2023 across 32 monitoring stations, we develop predictive regression models that incorporate meteorological variables (temperature and relative humidity), pollutant concentrations ($NO_2, SO_2, CO$), and seasonal harmonic components to capture both diurnal and annual cycles. Here, we show that reductions in anthropogenic emissions lead to significant decreases in $PM_{2.5}$ levels, whereas their effect on $O_3$ remains marginal and statistically insignificant. To address spatial heterogeneity, we employ Gaussian Process modeling. Further, we use Granger causality analysis and counterfactual simulation to establish direct causal links. Validation using real-world data from the COVID-19 lockdown confirms that reduced emissions led to a substantial drop in $PM_{2.5}$ but only a slight, insignificant change in $O_3$. The findings highlight the necessity of targeted emission reduction policies while emphasizing the need for integrated strategies addressing both particulate and ozone pollution. These insights are crucial for policymakers designing air pollution interventions in other megacities, and offer a scalable methodology for tackling complex urban air pollution through data-driven decision-making.

stat.AP

Understanding the Effect of Market Risks on New Pension System and Government Responsibility

This study examines how market risks impact the sustainability and performance of the New Pension System (NPS). NPS relies on defined contributions from both employees and employers to build a corpus during the employee's service period. Upon retirement, employees use the corpus fund to sustain their livelihood. A critical concern for individuals is whether the corpus will grow sufficiently to be sustainable or if it will deplete, leaving them financially vulnerable at an advanced age. We explore the impact of market risks on the performance of the corpus resulting from the NPS. To address this, we quantify market risks using Monte Carlo simulations with historical data to model their impact on NPS. We quantify the risk of pension corpus being insufficient and the cost to the Government to hedge the risk arising from guaranteeing the pension.

q-fin.RM

Arctic teleconnection on climate and ozone pollution in the polar jet stream path of eastern US

Arctic sea-ice loss is a defining feature of climate change and offers insight into its impact on mid-latitude air quality. Here, we investigate how variability in Arctic sea-ice extent (ASI) affects ground-level ozone ($O_3$) across eastern US states through physically and chemically mediated atmospheric pathways. Using observations and causal-inference methods grounded in atmospheric dynamics, we show that ASI drives wintertime ozone variability primarily via indirect meteorological mechanisms, including changes in humidity, temperature, and atmospheric circulation along the polar and subtropical jet streams. Inland regions exhibit the strongest sensitivity, while coastal areas are modulated by marine boundary-layer processes. Seasonal contrasts reveal that Arctic-driven dynamics suppress ozone in winter but can enhance accumulation under certain summer conditions. These findings highlight the importance of Arctic-midlatitude teleconnections in shaping regional air quality and highlight the need to integrate large-scale climate processes into ozone management and climate adaptation strategies.

physics.ao-ph

Understanding North Atlantic Climate Instabilities and Complex Interactions using Data Science

The North Atlantic Oscillation (NAO) index, a measure of sea-level atmospheric pressure variability, holds significant influence over weather patterns in North America and Northern Europe. A negative (positive) NAO value signifies increased cold air outbreaks and storm occurrences (reduced occurrences) in these regions. NAO, a product of multiple climate factors, demonstrates intricate dynamics with sea surface temperature (SST) and sea ice extent (SIE). In this study, we adopt a data-driven approach to explore the complex interplay between NAO, SST, and SIE, revealing a critical instability rooted in positive feedback loops among these climate variables. Our statistical machine learning methodology examines the impacts of melting Arctic SIE and rising SST on NAO, thereby understanding the weather patterns across the North Atlantic region. The skewness analysis yields a negative skewness in NAO across various time intervals -- daily, weekly, and monthly. This skewness, coupled with NAO's mean zero stationary nature, accentuates system instability. To capture these dynamics, we formulate a Bayesian Granger-causal dynamic linear model, which effectively updates the predictor-dependent variable relationship over time. The findings underscore an impending critical instability, indicative of more frequent occurrences of intensely cold climates in eastern North America and northern Europe, theory signifies a notable climate shift. By delving into the intricate feedback mechanisms of NAO, SST, and SIE, our study enhances our comprehension of climate variability, fostering a more informed perspective on the imminent climate changes that lie ahead.

stat.AP

Risk Analysis of Passive Portfolios

In this work, we present an alternative passive investment strategy. The passive investment philosophy comes from the Efficient Market Hypothesis (EMH), and its adoption is widespread. If EMH is true, one cannot outperform market by actively managing their portfolio for a long time. Also, it requires little to no intervention. People can buy an exchange-traded fund (ETF) with a long-term perspective. As the economy grows over time, one expects the ETF to grow. For example, in India, one can invest in NETF, which suppose to mimic the Nifty50 return. However, the weights of the Nifty 50 index are based on market capitalisation. These weights are not necessarily optimal for the investor. In this work, we present that volatility risk and extreme risk measures of the Nifty50 portfolio are uniformly larger than Markowitz's optimal portfolio. However, common people can't create an optimised portfolio. So we proposed an alternative passive investment strategy of an equal-weight portfolio. We show that if one pushes the maximum weight of the portfolio towards equal weight, the idiosyncratic risk of the portfolio would be minimal. The empirical evidence indicates that the risk profile of an equal-weight portfolio is similar to that of Markowitz's optimal portfolio. Hence instead of buying Nifty50 ETFs, one should equally invest in the stocks of Nifty50 to achieve a uniformly better risk profile than the Nifty 50 ETF portfolio. We also present an analysis of how portfolios perform to idiosyncratic events like the Russian invasion of Ukraine. We found that the equal weight portfolio has a uniformly lower risk than the Nifty 50 portfolio before and during the Russia-Ukraine war. All codes are available on GitHub (\url{https://github.com/sourish-cmi/quant/tree/main/Chap_Risk_Anal_of_Passive_Portfolio}).

q-fin.ST

Untangling Climate's Complexity: Methodological Insights

In this article, we review the interdisciplinary techniques (borrowed from physics, mathematics, statistics, machine-learning, etc.) and methodological framework that we have used to understand climate systems, which serve as examples of "complex systems". We believe that this would offer valuable insights to comprehend the complexity of climate variability and pave the way for drafting policies for action against climate change, etc. Our basic aim is to analyse time-series data structures across diverse climate parameters, extract Fourier-transformed features to recognize and model the trends/seasonalities in the climate variables using standard methods like detrended residual series analyses, correlation structures among climate parameters, Granger causal models, and other statistical machine-learning techniques. We cite and briefly explain two case studies: (i) the relationship between the Standardised Precipitation Index (SPI) and specific climate variables including Sea Surface Temperature (SST), El Niño Southern Oscillation (ENSO), and Indian Ocean Dipole (IOD), uncovering temporal shifts in correlations between SPI and these variables, and reveal complex patterns that drive drought and wet climate conditions in South-West Australia; (ii) the complex interactions of North Atlantic Oscillation (NAO) index, with SST and sea ice extent (SIE), potentially arising from positive feedback loops.

physics.data-an

Understanding the complex dynamics of climate change in south-west Australia using Machine Learning

The Standardized Precipitation Index (SPI) is used to indicate the meteorological drought situation - a negative (or positive) value of SPI would imply a dry (or wet) condition in a region over a period. The climate system is an excellent example of a complex system since there is an interplay and inter-relation of several climate variables. It is not always easy to identify the factors that may influence the SPI, or their inter-relations (including feedback loops). Here, we aim to study the complex dynamics that SPI has with the SST, NINO 3.4 and Indian Ocean Dipole (IOD), using a machine learning approach. Our findings are: (i) IOD was negatively correlated to SPI till 2008; (ii) until 2004, SST was negatively correlated with SPI; (iii) from 2005 to 2014, the SST had swung between negative and positive correlations; (iv) since 2014, we observed that the regression coefficient ($δ$) corresponding to SST has always been positive; (v) the SST has an upward trend, and the positive upward trend of $δ$ implied that SPI has been positively correlated with SST in recent years; and finally, (vi) the current value of SPI has a significant positive correlation with a past SPI value with a periodicity of about 7.5 years. Examining the complex dynamics, we used a statistical machine learning approach to construct an inferential network of these climate variables, which revealed that SST and NINO 3.4 directly couples with SPI, whereas IOD indirectly couples with SPI through SST and NINO 3.4. The system also indicated that Nino 3.4 has a significant negative effect on SPI. Interestingly, there seems to be a structural change in the complex dynamics of the four climate variables, some time in 2008. Though a simple 12-month moving average of SPI has a negative trend towards drought, the complex dynamics of SPI with other climate variables indicate a wet season for western Australia.

physics.data-an

Sparse Portfolio selection via Bayesian Multiple testing

We presented Bayesian portfolio selection strategy, via the $k$ factor asset pricing model. If the market is information efficient, the proposed strategy will mimic the market; otherwise, the strategy will outperform the market. The strategy depends on the selection of a portfolio via Bayesian multiple testing methodologies. We present the "discrete-mixture prior" model and the "hierarchical Bayes model with horseshoe prior." We define the Oracle set and prove that asymptotically the Bayes rule attains the risk of Bayes Oracle up to $O(1)$. Our proposed Bayes Oracle test guarantees statistical power by providing the upper bound of the type-II error. Simulation study indicates that the proposed Bayes oracle test is suitable for the efficient market with few stocks inefficiently priced. However, as the model becomes dense, i.e., the market is highly inefficient, one should not use the Bayes oracle test. The statistical power of the Bayes Oracle portfolio is uniformly better for the $k$-factor model ($k>1$) than the one factor CAPM. We present the empirical study, where we considered the 500 constituent stocks of S\&P 500 from the New York Stock Exchange (NYSE), and S\&P 500 index as the benchmark for thirteen years from the year 2006 to 2018. We showed the out-sample risk and return performance of the four different portfolio selection strategies and compared with the S\&P 500 index as the benchmark market index. Empirical results indicate that it is possible to propose a strategy which can outperform the market.

q-fin.MF

Prediction of COVID-19 Disease Progression in India : Under the Effect of National Lockdown

In this policy paper, we implement the epidemiological SIR to estimate the basic reproduction number $\mathcal{R}_0$ at national and state level. We also developed the statistical machine learning model to predict the cases ahead of time. Our analysis indicates that the situation of Punjab ($\mathcal{R}_0\approx 16$) is not good. It requires immediate aggressive attention. We see the $\mathcal{R}_0$ for Madhya Pradesh (3.37) , Maharastra (3.25) and Tamil Nadu (3.09) are more than 3. The $\mathcal{R}_0$ of Andhra Pradesh (2.96), Delhi (2.82) and West Bengal (2.77) is more than the India's $\mathcal{R}_0=2.75$, as of 04 March, 2020. India's $\mathcal{R}_0=2.75$ (as of 04 March, 2020) is very much comparable to Hubei/China at the early disease progression stage. Our analysis indicates that the early disease progression of India is that of similar to China. Therefore, with lockdown in place, India should expect as many as cases if not more like China. If lockdown works, we should expect less than 66,224 cases by May 01,2020. All data and \texttt{R} code for this paper is available from \url{https://github.com/sourish-cmi/Covid19}

q-bio.PE

Causal Impact of Web Browsing and Other Factors on Research Publications

In this paper, we study the causal impact of the web-search activity on the research publication. We considered observational prospective study design, where research activity of 267 scientists is being studied. We considered the Poisson and negative binomial regression model for our analysis. Based on the Akaike's Model selection criterion, we found the negative binomial regression performs better than the Poisson regression. Detailed analysis indicates that the higher web-search activity of 2016 related to the sci-indexed website has a positive significant impact on the research publication of 2017. We observed that unique collaborations of 2016 and web-search activity of 2016 have a non-linear but significant positive impact on the research publication of 2017. What-if analysis indicates the high web browsing activity leads to more number of the publication. However, interestingly we see a scientist with low web activity can be as productive as others if her/his maximum hits are the sci-indexed journal. That is if the scientist uses web browsing only for research-related activity, then she/he can be equally productive even if her/his web activity is lower than fellow scientists.

stat.AP

A Bayesian Perspective of Statistical Machine Learning for Big Data

Statistical Machine Learning (SML) refers to a body of algorithms and methods by which computers are allowed to discover important features of input data sets which are often very large in size. The very task of feature discovery from data is essentially the meaning of the keyword `learning' in SML. Theoretical justifications for the effectiveness of the SML algorithms are underpinned by sound principles from different disciplines, such as Computer Science and Statistics. The theoretical underpinnings particularly justified by statistical inference methods are together termed as statistical learning theory. This paper provides a review of SML from a Bayesian decision theoretic point of view -- where we argue that many SML techniques are closely connected to making inference by using the so called Bayesian paradigm. We discuss many important SML techniques such as supervised and unsupervised learning, deep learning, online learning and Gaussian processes especially in the context of very large data sets where these are often employed. We present a dictionary which maps the key concepts of SML from Computer Science and Statistics. We illustrate the SML techniques with three moderately large data sets where we also discuss many practical implementation issues. Thus the review is especially targeted at statisticians and computer scientists who are aspiring to understand and apply SML for moderately large to big data sets.

cs.LG

Modeling Nelson-Siegel Yield Curve using Bayesian Approach

Yield curve modeling is an essential problem in finance. In this work, we explore the use of Bayesian statistical methods in conjunction with Nelson-Siegel model. We present the hierarchical Bayesian model for the parameters of the Nelson-Siegel yield function. We implement the MAP estimates via BFGS algorithm in rstan. The Bayesian analysis relies on the Monte Carlo simulation method. We perform the Hamiltonian Monte Carlo (HMC), using the rstan package. As a by-product of the HMC, we can simulate the Monte Carlo price of a Bond, and it helps us to identify if the bond is over-valued or under-valued. We demonstrate the process with an experiment and US Treasury's yield curve data. One of the interesting observation of the experiment is that there is a strong negative correlation between the price and long-term effect of yield. However, the relationship between the short-term interest rate effect and the value of the bond is weakly positive. This is because posterior analysis shows that the short-term effect and the long-term effect are negatively correlated.

q-fin.ST

Modeling Risk and Return using Dirichlet Process Prior

In this paper, we showed that the no-arbitrage condition holds if the market follows the mixture of the geometric Brownian motion (GBM). The mixture of GBM can incorporate heavy-tail behavior of the market. It automatically leads us to model the risk and return of multiple asset portfolios via the nonparametric Bayesian method. We present a Dirichlet Process (DP) prior via an urn-scheme for univariate modeling of the single asset return. This DP prior is presented in the spirit of dependent DP. We extend this approach to introduce a multivariate distribution to model the return on multiple assets via an elliptical copula; which models the marginal distribution using the DP prior. We compare different risk measures such as Value at Risk (VaR) and Conditional VaR (CVaR), also known as expected shortfall (ES) for the stock return data of two datasets. The first dataset contains the return of IBM, Intel and NASDAQ and the second dataset contains the return data of 51 stocks as part of the index "Nifty 50" for Indian equity markets.

stat.ME