SearcharxivSearch

arXiv subjects

Sujit K. Ghosh

Publications and source records attributed to Sujit K. Ghosh.

At least 19 recordsLinked to original sources

Nonparametric Estimation of Isotropic Covariance Function

A nonparametric model using a sequence of Bernstein polynomials is constructed to approximate arbitrary isotropic covariance functions valid in $\mathbb{R}^\infty$ and related approximation properties are investigated using the popular $L_{\infty}$ norm and $L_2$ norms. A computationally efficient sieve maximum likelihood (sML) estimation is then developed to nonparametrically estimate the unknown isotropic covaraince function valid in $\mathbb{R}^\infty$. Consistency of the proposed sieve ML estimator is established under increasing domain regime. The proposed methodology is compared numerically with couple of existing nonparametric as well as with commonly used parametric methods. Numerical results based on simulated data show that our approach outperforms the parametric methods in reducing bias due to model misspecification and also the nonparametric methods in terms of having significantly lower values of expected $L_{\infty}$ and $L_2$ norms. Application to precipitation data is illustrated to showcase a real case study. Additional technical details and numerical illustrations are also made available.

stat.ME

Forecasting Malaria in Indian States: A Time Series Approach with R Shiny Integration

Malaria remains a significant public health challenge in many regions, necessitating robust predictive models to aid in its management and prevention. This study focuses on developing and evaluating time series models for forecasting malaria cases across eight Indian states: Jharkhand, Chhattisgarh, Maharashtra, Meghalaya, Mizoram, Odisha, Tripura, and Uttar Pradesh. We employed various modeling approaches, including polynomial regression with seasonal components, log-transformed polynomial regression, lagged difference models, and ARIMA models, to capture the temporal dynamics of malaria incidence. Comprehensive model fitting, residual analysis, and performance evaluation using metrics such as Root Mean Squared Error (RMSE) and Mean Absolute Percentage Error (MAPE) indicated that the log-transformed polynomial regression model consistently outperformed other models in terms of accuracy and robustness across all states. Rolling forecast validation further confirmed the superior predictive capability of the log-transformed model over time. Additionally, an interactive R Shiny tool was developed to facilitate the use of these predictive models by researchers and public health officials. This tool allows users to input data, select modeling approaches, and visualize predictions and performance metrics, providing a practical tool for real-time malaria forecasting and decision-making support. Our findings highlight the critical role of appropriate modeling techniques in malaria prediction and offer valuable resources for enhancing malaria surveillance and response efforts.

stat.AP

Bayesian Models for Joint Selection of Features and Auto-Regressive Lags: Theory and Applications in Environmental and Financial Forecasting

We develop a Bayesian framework for variable selection in linear regression with autocorrelated errors, accommodating lagged covariates and autoregressive structures. This setting occurs in time series applications where responses depend on contemporaneous or past explanatory variables and persistent stochastic shocks, including financial modeling, hydrological forecasting, and meteorological applications requiring temporal dependency capture. Our methodology uses hierarchical Bayesian models with spike-and-slab priors to simultaneously select relevant covariates and lagged error terms. We propose an efficient two-stage MCMC algorithm separating sampling of variable inclusion indicators and model parameters to address high-dimensional computational challenges. Theoretical analysis establishes posterior selection consistency under mild conditions, even when candidate predictors grow exponentially with sample size, common in modern time series with many potential lagged variables. Through simulations and real applications (groundwater depth prediction, S&P 500 log returns modeling), we demonstrate substantial gains in variable selection accuracy and predictive performance. Compared to existing methods, our framework achieves lower MSPE, improved true model component identification, and greater robustness with autocorrelated noise, underscoring practical utility for model interpretation and forecasting in autoregressive settings.

stat.ME

Semiparametric Dynamic Copula Models for Portfolio Optimization

The mean-variance portfolio model, based on the risk-return trade-off for optimal asset allocation, remains foundational in portfolio optimization. However, its reliance on restrictive assumptions about asset return distributions limits its applicability to real-world data. Parametric copula structures provide a novel way to overcome these limitations by accounting for asymmetry, heavy tails, and time-varying dependencies. Existing methods have been shown to rely on fixed or static dependence structures, thus overlooking the dynamic nature of the financial market. In this study, a semiparametric model is proposed that combines non-parametrically estimated copulas with parametrically estimated marginals to allow all parameters to dynamically evolve over time. A novel framework was developed that integrates time-varying dependence modeling with flexible empirical beta copula structures. Marginal distributions were modeled using the Skewed Generalized T family. This effectively captures asymmetry and heavy tails and makes the model suitable for predictive inferences in real world scenarios. Furthermore, the model was applied to rolling windows of financial returns from the USA, India and Hong Kong economies to understand the influence of dynamic market conditions. The approach addresses the limitations of models that rely on parametric assumptions. By accounting for asymmetry, heavy tails, and cross-correlated asset prices, the proposed method offers a robust solution for optimizing diverse portfolios in an interconnected financial market. Through adaptive modeling, it allows for better management of risk and return across varying economic conditions, leading to more efficient asset allocation and improved portfolio performance.

q-fin.PM

Bayesian Modal Regression for Forecast Combinations

Forecast combination methods have traditionally emphasized symmetric loss functions, particularly squared error loss, with equally weighted combinations often justified as a robust approach under such criteria. However, these justifications do not extend to asymmetric loss functions, where optimally weighted combinations may provide superior predictive performance. This study introduces a novel contribution by incorporating modal regression into forecast combinations, offering a Bayesian hierarchical framework that models the conditional mode of the response through combinations of time-varying parameters and exponential discounting. The proposed approach utilizes error distributions characterized by asymmetry and heavy tails, specifically the asymmetric Laplace, asymmetric normal, and reverse Gumbel distributions. Simulated data validate the parameter estimation for the modal regression models, confirming the robustness of the proposed methodology. Application of these methodologies to a real-world analyst forecast dataset shows that modal regression with asymmetric Laplace errors outperforms mean regression based on two key performance metrics: the hit rate, which measures the accuracy of classifying the sign of revenue surprises, and the win rate, which assesses the proportion of forecasts surpassing the equally weighted consensus. These results underscore the presence of skewness and fat-tailed behavior in forecast combination errors for revenue forecasting, highlighting the advantages of modal regression in financial applications.

stat.ME

Optimizing Forecast Combination Weights Using Exponentially Weighted Hit and Win Rate Losses

Forecasting revenues by aggregating analyst forecasts is a fundamental problem in financial research and practice. A key objective in this context is to improve the accuracy of the forecast by optimizing two performance metrics: the hit rate, which measures the proportion of correctly classified revenue surprise signs, and the win rate, which quantifies the proportion of individual forecasts that outperform an equally weighted consensus benchmark. While researchers have extensively studied forecast combination techniques, two critical gaps remain: (i) the estimation of optimal combination weights tailored to these specific performance metrics and (ii) the development of Bayesian methods for handling missing or incomplete analyst forecasts. This paper proposes novel approaches to address these challenges. First, we introduce a method for estimating optimal forecast combination weights using exponentially weighted hit and win rate loss functions via nonlinear programming. Second, we develop a Bayesian imputation framework that leverages exponentially weighted likelihood methods to account for missing forecasts while preserving key distributional properties. Through extensive empirical evaluations using real-world analyst forecast data, we demonstrate that our proposed methodologies yield superior predictive performance compared to traditional equally weighted and linear combination benchmarks. These findings highlight the advantages of incorporating tailored loss functions and Bayesian inference in forecast combination models, offering valuable insights for financial analysts and practitioners seeking to improve revenue prediction accuracy.

stat.ME

A new block covariance regression model and inferential framework for massively large neuroimaging data

Some evidence suggests that people with autism spectrum disorder exhibit patterns of brain functional dysconnectivity relative to their typically developing peers, but specific findings have yet to be replicated. To facilitate this replication goal with data from the Autism Brain Imaging Data Exchange (ABIDE), we propose a flexible and interpretable model for participant-specific voxel-level brain functional connectivity. Our approach efficiently handles massive participant-specific whole brain voxel-level connectivity data that exceed one trillion data points. The key component of the model is to leverage the block structure induced by defined regions of interest to introduce parsimony in the high-dimensional connectivity matrix through a block covariance structure. Associations between brain functional connectivity and participant characteristics -- including eye status during the resting scan, sex, age, and their interactions -- are estimated within a Bayesian framework. A spike-and-slab prior facilitates hypothesis testing to identify voxels associated with autism diagnosis. Simulation studies are conducted to evaluate the empirical performance of the proposed model and estimation framework. In ABIDE, the method replicates key findings from the literature and suggests new associations for investigation.

stat.ME

Functional Time Transformation Model with Applications to Digital Health

The advent of wearable and sensor technologies now leads to functional predictors which are intrinsically infinite dimensional. While the existing approaches for functional data and survival outcomes lean on the well-established Cox model, the proportional hazard (PH) assumption might not always be suitable in real-world applications. Motivated by physiological signals encountered in digital medicine, we develop a more general and flexible functional time-transformation model for estimating the conditional survival function with both functional and scalar covariates. A partially functional regression model is used to directly model the survival time on the covariates through an unknown monotone transformation and a known error distribution. We use Bernstein polynomials to model the monotone transformation function and the smooth functional coefficients. A sieve method of maximum likelihood is employed for estimation. Numerical simulations illustrate a satisfactory performance of the proposed method in estimation and inference. We demonstrate the application of the proposed model through two case studies involving wearable data i) Understanding the association between diurnal physical activity pattern and all-cause mortality based on accelerometer data from the National Health and Nutrition Examination Survey (NHANES) 2011-2014 and ii) Modelling Time-to-Hypoglycemia events in a cohort of diabetic patients based on distributional representation of continuous glucose monitoring (CGM) data. The results provide important epidemiological insights into the direct association between survival times and the physiological signals and also exhibit superior predictive performance compared to traditional summary based biomarkers in the CGM study.

stat.ME

Distributional outcome regression via quantile functions and its application to modelling continuously monitored heart rate and physical activity

Modern clinical and epidemiological studies widely employ wearables to record parallel streams of real-time data on human physiology and behavior. With recent advances in distributional data analysis, these high-frequency data are now often treated as distributional observations resulting in novel regression settings. Motivated by these modelling setups, we develop a distributional outcome regression via quantile functions (DORQF) that expands existing literature with three key contributions: i) handling both scalar and distributional predictors, ii) ensuring jointly monotone regression structure without enforcing monotonicity on individual functional regression coefficients, iii) providing statistical inference via asymptotic projection-based joint confidence bands and a statistical test of global significance to quantify uncertainty of the estimated functional regression coefficients. The method is motivated by and applied to Actiheart component of Baltimore Longitudinal Study of Aging that collected one week of minute-level heart rate (HR) and physical activity (PA) data on 781 older adults to gain deeper understanding of age-related changes in daily life heart rate reserve, defined as a distribution of daily HR, while accounting for daily distribution of physical activity, age, gender, and body composition. Intriguingly, the results provide novel insights in epidemiology of daily life heart rate reserve.

stat.ME

Bayesian estimation of clustered dependence structures in functional neuroconnectivity

Motivated by the need to model the dependence between regions of interest in functional neuroconnectivity for efficient inference, we propose a new sampling-based Bayesian clustering approach for covariance structures of high-dimensional Gaussian outcomes. The key technique is based on a Dirichlet process that clusters covariance sub-matrices into independent groups of outcomes, thereby naturally inducing sparsity in the whole brain connectivity matrix. A new split-merge algorithm is employed to achieve convergence of the Markov chain that is shown empirically to recover both uniform and Dirichlet partitions with high accuracy. We investigate the empirical performance of the proposed method through extensive simulations. Finally, the proposed approach is used to group regions of interest into functionally independent groups in the Autism Brain Imaging Data Exchange participants with autism spectrum disorder and and co-occurring attention-deficit/hyperactivity disorder.

stat.ME

Beyond 2-D Mass-Radius Relationships: A Nonparametric and Probabilistic Framework for Characterizing Planetary Samples in Higher Dimensions

Fundamental to our understanding of planetary bulk compositions is the relationship between their masses and radii, two properties that are often not simultaneously known for most exoplanets. However, while many previous studies have modeled the two-dimensional relationship between planetary mass and radii, this approach largely ignores the dependencies on other properties that may have influenced the formation and evolution of the planets. In this work, we extend the existing nonparametric and probabilistic framework of \texttt{MRExo} to jointly model distributions beyond two dimensions. Our updated framework can now simultaneously model up to four observables, while also incorporating asymmetric measurement uncertainties and upper limits in the data. We showcase the potential of this multi-dimensional approach to three science cases: (i) a 4-dimensional joint fit to planetary mass, radius, insolation, and stellar mass, hinting of changes in planetary bulk density across insolation and stellar mass; (ii) a 3-dimensional fit to the California Kepler Survey sample showing how the planet radius valley evolves across different stellar masses; and (iii) a 2-dimensional fit to a sample of Class-II protoplanetary disks in Lupus while incorporating the upper-limits in dust mass measurements. In addition, we employ bootstrap and Monte-Carlo sampling to quantify the impact of the finite sample size as well as measurement uncertainties on the predicted quantities. We update our existing open-source user-friendly \texttt{MRExo} \texttt{Python} package with these changes, which allows users to apply this highly flexible framework to a variety of datasets beyond what we have shown here.

astro-ph.EP

EMFlow: Data Imputation in Latent Space via EM and Deep Flow Models

The presence of missing values within high-dimensional data is an ubiquitous problem for many applied sciences. A serious limitation of many available data mining and machine learning methods is their inability to handle partially missing values and so an integrated approach that combines imputation and model estimation is vital for down-stream analysis. A computationally fast algorithm, called EMFlow, is introduced that performs imputation in a latent space via an online version of Expectation-Maximization (EM) algorithm by using a normalizing flow (NF) model which maps the data space to a latent space. The proposed EMFlow algorithm is iterative, involving updating the parameters of online EM and NF alternatively. Extensive experimental results for high-dimensional multivariate and image datasets are presented to illustrate the superior performance of the EMFlow compared to a couple of recently available methods in terms of both predictive accuracy and speed of algorithmic convergence. We provide code for all our experiments.

cs.LG

Bayesian Inference for Generalized Linear Model with Linear Inequality Constraints

Bayesian statistical inference for Generalized Linear Models (GLMs) with parameters lying on a constrained space is of general interest (e.g., in monotonic or convex regression), but often constructing valid prior distributions supported on a subspace spanned by a set of linear inequality constraints can be challenging, especially when some of the constraints might be binding leading to a lower dimensional subspace. For the general case with canonical link, it is shown that a generalized truncated multivariate normal supported on a desired subspace can be used. Moreover, it is shown that such prior distribution facilitates the construction of a general purpose product slice sampling method to obtain (approximate) samples from corresponding posterior distribution, making the inferential method computationally efficient for a wide class of GLMs with an arbitrary set of linear inequality constraints. The proposed product slice sampler is shown to be uniformly ergodic, having a geometric convergence rate under a set of mild regularity conditions satisfied by many popular GLMs (e.g., logistic and Poisson regressions with constrained coefficients). One of the primary advantages of the proposed Bayesian estimation method over classical methods is that uncertainty of parameter estimates is easily quantified by using the samples simulated from the path of the Markov Chain of the slice sampler. Numerical illustrations using simulated data sets are presented to illustrate the superiority of the proposed methods compared to some existing methods in terms of sampling bias and variances. In addition, real case studies are presented using data sets for fertilizer-crop production and estimating the SCRAM rate in nuclear power plants.

stat.ME

A Non-Iterative Quantile Change Detection Method in Mixture Model with Heavy-Tailed Components

Estimating parameters of mixture model has wide applications ranging from classification problems to estimating of complex distributions. Most of the current literature on estimating the parameters of the mixture densities are based on iterative Expectation Maximization (EM) type algorithms which require the use of either taking expectations over the latent label variables or generating samples from the conditional distribution of such latent labels using the Bayes rule. Moreover, when the number of components is unknown, the problem becomes computationally more demanding due to well-known label switching issues \cite{richardson1997bayesian}. In this paper, we propose a robust and quick approach based on change-point methods to determine the number of mixture components that works for almost any location-scale families even when the components are heavy tailed (e.g., Cauchy). We present several numerical illustrations by comparing our method with some of popular methods available in the literature using simulated data and real case studies. The proposed method is shown be as much as 500 times faster than some of the competing methods and are also shown to be more accurate in estimating the mixture distributions by goodness-of-fit tests.

stat.ML

Probabilistic Detection and Estimation of Conic Sections from Noisy Data

Inferring unknown conic sections on the basis of noisy data is a challenging problem with applications in computer vision. A major limitation of the currently available methods for conic sections is that estimation methods rely on the underlying shape of the conics (being known to be ellipse, parabola or hyperbola). A general purpose Bayesian hierarchical model is proposed for conic sections and corresponding estimation method based on noisy data is shown to work even when the specific nature of the conic section is unknown. The model, thus, provides probabilistic detection of the underlying conic section and inference about the associated parameters of the conic section. Through extensive simulation studies where the true conics may not be known, the methodology is demonstrated to have practical and methodological advantages relative to many existing techniques. In addition, the proposed method provides probabilistic measures of uncertainty of the estimated parameters. Furthermore, we observe high fidelity to the true conics even in challenging situations, such as data arising from partial conics in arbitrarily rotated and non-standard form, and where a visual inspection is unable to correctly identify the type of conic section underlying the data.

stat.ME

Data transforming augmentation for heteroscedastic models

Data augmentation (DA) turns seemingly intractable computational problems into simple ones by augmenting latent missing data. In addition to computational simplicity, it is now well-established that DA equipped with a deterministic transformation can improve the convergence speed of iterative algorithms such as an EM algorithm or Gibbs sampler. In this article, we outline a framework for the transformation-based DA, which we call data transforming augmentation (DTA), allowing augmented data to be a deterministic function of latent and observed data, and unknown parameters. Under this framework, we investigate a novel DTA scheme that turns heteroscedastic models into homoscedastic ones to take advantage of simpler computations typically available in homoscedastic cases. Applying this DTA scheme to fitting linear mixed models, we demonstrate simpler computations and faster convergence rates of resulting iterative algorithms, compared with those under a non-transformation-based DA scheme. We also fit a Beta-Binomial model using the proposed DTA scheme, which enables sampling approximate marginal posterior distributions that are available only under homoscedasticity. An R package, Rdta, is publicly available at CRAN.

stat.ME

An Unified Semiparametric Approach to Model Lifetime Data with Crossing Survival Curves

The proportional hazards (PH), proportional odds (PO) and accelerated failure time (AFT) models have been widely used in different applications of survival analysis. Despite their popularity, these models are not suitable to handle lifetime data with crossing survival curves. In 2005, Yang and Prentice proposed a semiparametric two-sample strategy (YP model), including the PH and PO frameworks as particular cases, to deal with this type of data. Assuming a general regression setting, the present paper proposes an unified approach to fit the YP model by employing Bernstein polynomials to manage the baseline hazard and odds under both the frequentist and Bayesian frameworks. The use of the Bernstein polynomials has some advantages: it allows for uniform approximation of the baseline distribution, it leads to closed-form expressions for all baseline functions, it simplifies the inference procedure, and the presence of a continuous survival function allows a more accurate estimation of the crossing survival time. Extensive simulation studies are carried out to evaluate the behavior of the models. The analysis of a clinical trial data set, related to non-small-cell lung cancer, is also developed as an illustration. Our findings indicate that assuming the usual PH model, ignoring the existing crossing survival feature in the real data, is a serious mistake with implications for those patients in the initial stage of treatment.

stat.ME

Assessing Biosimilarity using Functional Metrics

In recent years there have been a lot of interest to test for similarity between biological drug products, commonly known as biologics. Biologics are large and complex molecule drugs that are produced by living cells and hence these are sensitive to the environmental changes. In addition, biologics usually induce antibodies which raises the safety and efficacy issues. The manufacturing process is also much more complicated and often costlier than the small-molecule generic drugs. Because of these complexities and inherent variability of the biologics, the testing paradigm of the traditional generic drugs cannot be directly used to test for biosimilarity. Taking into account some of these concerns we propose a functional distance based methodology that takes into consideration the entire time course of the study and is based on a class of flexible semi-parametric models. The empirical results show that the proposed approach is more sensitive than the classical equivalence tests approach which are usually based on arbitrarily chosen time point. Bootstrap based methodologies are also presented for statistical inference.

stat.ME