SearcharxivSearch

arXiv subjects

Raju Maiti

Publications and source records attributed to Raju Maiti.

8 recordsLinked to original sources

Forecasting of Multiple Seasonal Categorical Time Series Using Fourier Series with Application to AQI Data of Kolkata

Multiple seasonalities have been widely studied in continuous time series using models such as TBATS, for instance in electricity demand forecasting. However, their treatment in categorical time series, such as air quality index (AQI) data, remains limited. Categorical AQI often exhibits distinct seasonal patterns at multiple frequencies, which are not captured by standard models. In this paper, we propose a framework that models multiple seasonalities using Fourier series and indicator functions, inspired by the TBATS methodology. The approach accommodates the ordinal nature of AQI categories while explicitly capturing daily, weekly and yearly seasonal cycles. Simulation studies demonstrate the empirical consistency of parameter estimates under the proposed model. We further illustrate its applicability using real categorical AQI data from Kolkata and compare forecasting performance with Markov models and machine learning methods. Results indicate that our approach effectively captures complex seasonal dynamics and provides improved predictive accuracy. The proposed methodology offers a flexible and interpretable framework for analyzing categorical time series exhibiting multiple seasonal patterns, with potential applications in air quality monitoring, energy consumption and other environmental domains.

stat.ME

Long-Term Spatio-Temporal Forecasting of Monthly Rainfall in West Bengal Using Ensemble Learning Approaches

Rainfall forecasting plays a critical role in climate adaptation, agriculture, and water resource management. This study develops long-term forecasts of monthly rainfall across 19 districts of West Bengal using a century-scale dataset spanning 1900-2019. Daily rainfall records are aggregated into monthly series, resulting in 120 years of observations for each district. The forecasting task involves predicting the next 108 months (9 years, 2011-2019) while accounting for temporal dependencies and spatial interactions among districts. To address the nonlinear and complex structure of rainfall dynamics, we propose a hierarchical modeling framework that combines regression-based forecasting of yearly features with multi-layer perceptrons (MLPs) for monthly prediction. Yearly features, such as annual totals, quarterly proportions, variability measures, skewness, and extremes, are first forecasted using regression models that incorporate both own lags and neighboring-district lags. These forecasts are then integrated as auxiliary inputs into an MLP model, which captures nonlinear temporal patterns and spatial dependencies in the monthly series. The results demonstrate that the hierarchical regression-MLP architecture provides robust long-term spatio-temporal forecasts, offering valuable insights for agriculture, irrigation planning, and water conservation strategies.

stat.AP

Sample size estimation for comparing dynamic treatment regimens in a SMART: a Monte Carlo-based approach and case study with longitudinal overdispersed count outcomes

Dynamic treatment regimens (DTRs), also known as treatment algorithms or adaptive interventions, play an increasingly important role in many health domains. DTRs are motivated to address the unique and changing needs of individuals by delivering the type of treatment needed, when needed, while minimizing unnecessary treatment. Practically, a DTR is a sequence of decision rules that specify, for each of several points in time, how available information about the individual's status and progress should be used in practice to decide which treatment (e.g., type or intensity) to deliver. The sequential multiple assignment randomized trial (SMART) is an experimental design widely used to empirically inform the development of DTRs. Sample size planning resources for SMARTs have been developed for continuous, binary, and survival outcomes. However, an important gap exists in sample size estimation methodology for SMARTs with longitudinal count outcomes. Further, in many health domains, count data are overdispersed - having variance greater than their mean. We propose a Monte Carlo-based approach to sample size estimation applicable to many types of longitudinal outcomes and provide a case study with longitudinal overdispersed count outcomes. A SMART for engaging alcohol and cocaine-dependent patients in treatment is used as motivation.

stat.ME

Scaling Behavior of the Hirsch Index for Failure Avalanches, Percolation Clusters and Paper Citations

A popular measure for citation inequalities of individual scientists has been the Hirsch index ($h$). If for any scientist the number $n_c$ of citations is plotted against the serial number $n_p$ of the paper having those many citations (when the papers are ordered from highest cited to lowest) then $h$ corresponds to the nearest lower integer value of $n_p$ below the fixed point of the non-linear citation function (or given by $n_c = h = n_p$ if both $n_p$ and $n_c$ are dense set of integers near the $h$ value). The same index can be estimated (from $h=s=n_{s}$) for the avalanche or cluster of size ($s$) distributions ($n_s$) in elastic fiber bundle or percolation models. Another such inequality index, called the Kolkata index ($k$) says that $(1-k)$ fraction of papers attract $k$ fraction of citations ($k=0.80$ corresponds to the 80-20 law of Pareto). We find, for stress ($σ$), lattice occupation probability ($p$) or Kolkata index ($k$) near the bundle failure threshold ($σ_c$) or percolation threshold ($p_c$) or critical value of Kolkata index $k_c$, good fit to Widom-Stauffer like scaling $h/[\sqrt{N}/log N]$ = $f(\sqrt{N}[σ_c -σ]^α)$, $h/[\sqrt{N}/log N]=f(\sqrt{N}|p_c -p|^α)$ or $h/[\sqrt{N_c}/log N_c]=f(\sqrt{N_c}|k_c -k|^α)$ respectively, with asymptotically defined scaling function $f$, for systems of size $N$ (total number of fibers or lattice sites) or $N_c$ (total number of citations), and $α$ denoting the appropriate scaling exponent. We also show that if the number ($N_m$) of members of parliaments or national assemblies of different countries (with population $N$) is identified as their respective $h-$index, then the data fit the scaling relation $N_m \sim \sqrt N /log N$, resolving a major recent controversy.

physics.soc-ph

Evolutionary Dynamics of Social Inequality and Coincidence of Gini and Kolkata indices under Unrestricted Competition

Social inequalities are ubiquitous and here we show that the values of the Gini ($g$) and Kolkata ($k$) indices, two generic inequality indices, approach each other (starting from $g = 0$ and $k = 0.5$ for equality) as the competitions grow in various social institutions like markets, universities, elections, etc. It is further showed that these two indices become equal and stabilize at a value (at $g = k \simeq 0.87$) under unrestricted competitions. We propose to view this coincidence of inequality indices as a generalized version of the (more than a) century old 80-20 law of Pareto. Furthermore, the coincidence of the inequality indices noted here is very similar to the ones seen before for self-organized critical (SOC) systems. The observations here, therefore, stand as a quantitative support towards viewing interacting socio-economic systems in the framework of SOC, an idea conjectured for years.

physics.soc-ph

Estimating the Optimal Linear Combination of Biomarkers using Spherically Constrained Optimization

In the context of a binary classification problem, the optimal linear combination of continuous predictors can be estimated by maximizing an empirical estimate of the area under the receiver operating characteristic (ROC) curve (AUC). For multi-category responses, the optimal predictor combination can similarly be obtained by maximization of the empirical hypervolume under the manifold (HUM). This problem is particularly relevant to medical research, where it may be of interest to diagnose a disease with various subtypes or predict a multi-category outcome. Since the empirical HUM is discontinuous, non-differentiable, and possibly multi-modal, solving this maximization problem requires a global optimization technique. Estimation of the optimal coefficient vector using existing global optimization techniques is computationally expensive, becoming prohibitive as the number of predictors and the number of outcome categories increases. We propose an efficient derivative-free black-box optimization technique based on pattern search to solve this problem. Through extensive simulation studies, we demonstrate that the proposed method achieves better performance compared to existing methods including the step-down algorithm. Finally, we illustrate the proposed method to predict swallowing difficulty after radiation therapy for oropharyngeal cancer based on radiation dose to various structures in the head and neck.

stat.ME

A Sequential Significance Test for Treatment by Covariate Interactions

Due to patient heterogeneity in response to various aspects of any treatment program, biomedical and clinical research is gradually shifting from the traditional "one-size-fits-all" approach to the new paradigm of personalized medicine. An important step in this direction is to identify the treatment by covariate interactions. We consider the setting in which there are potentially a large number of covariates of interest. Although a number of novel machine learning methodologies have been developed in recent years to aid in treatment selection in this setting, few, if any, have adopted formal hypothesis testing procedures. In this article, we present a novel testing procedure based on m-out-of-n bootstrap that can be used to sequentially identify variables that interact with treatment. We study the theoretical properties of the method and show that it is more effective in controlling the type I error rate and achieving a satisfactory power as compared to competing methods, via extensive simulations. Furthermore, the usefulness of the proposed method is illustrated using real data examples, both from a randomized trial and from an observational study.

stat.ME

A distribution-free smoothed combination method of biomarkers to improve diagnostic accuracy in multi-category classification

Results from multiple diagnostic tests are usually combined to improve the overall diagnostic accuracy. For binary classification, maximization of the empirical estimate of the area under the receiver operating characteristic (ROC) curve is widely adopted to produce the optimal linear combination of multiple biomarkers. In the presence of large number of biomarkers, this method proves to be computationally expensive and difficult to implement since it involves maximization of a discontinuous, non-smooth function for which gradient-based methods cannot be used directly. Complexity of this problem increases when the classification problem becomes multi-category. In this article, we develop a linear combination method that maximizes a smooth approximation of the empirical Hypervolume Under Manifolds (HUM) for multi-category outcome. We approximate HUM by replacing the indicator function with the sigmoid function or normal cumulative distribution function (CDF). With the above smooth approximations, efficient gradient-based algorithms can be employed to obtain better solution with less computing time. We show that under some regularity conditions, the proposed method yields consistent estimates of the coefficient parameters. We also derive the asymptotic normality of the coefficient estimates. We conduct extensive simulations to examine our methods. Under different simulation scenarios, the proposed methods are compared with other existing methods and are shown to outperform them in terms of diagnostic accuracy. The proposed method is illustrated using two real medical data sets.

stat.ME