SearcharxivSearch

arXiv subjects

Surojit Biswas

Publications and source records attributed to Surojit Biswas.

14 recordsLinked to original sources

Distribution-Free test for Changepoint Detection in Angular Mean Direction: Application in Finance

In this paper, we propose a distribution-free test for detecting changepoint in the mean direction of angular data. The uncertainty in angular measurements is quantified through the \textit{square of an angle}, derived from the intrinsic geometry of the torus. It is established that, under the null hypothesis, the test statistic distributionally converges to the Kolmogorov distribution, while under the alternative hypothesis, both the consistency of the test and the asymptotic properties of the changepoint estimator are established. Through extensive simulations, we compare the empirical performance of the proposed method with two existing approaches for angular data and further benchmark it against a test based on the circular arc length distance. Finally, we demonstrate the practical utility of our approach by analyzing the timestamps of extreme events in Bitcoin, Ethereum, and Gold price datasets, where the continuous, high-frequency nature of the data is modeled in the circular framework.

stat.ME

Intrinsic Geometry-Based Angular Covariance: A Novel Framework for Nonparametric Changepoint Detection in Meteorological Data

In many temporal datasets, the parameters of the underlying distribution may change abruptly at unknown times. Detecting such changepoints is crucial for numerous applications. Although such a problem has been extensively studied for linear data, there has been notably less research on bivariate angular data. To the best of our knowledge, this paper presents the first attempt to address the changepoint detection problem for the mean direction of toroidal and spherical data. By defining the ``square of an angle'' through intrinsic geometry, we construct a curved dispersion matrix for bivariate angular data, analogous to the linear dispersion matrix in Euclidean space. Using the analogous measure of the ``Mahalanobis distance,'' we develop two new non-parametric tests to identify changes in the mean direction parameters for toroidal and spherical distributions. The pivotal distributions of the test statistics are shown to follow the Kolmogorov distribution under the null hypothesis. Under the alternative hypothesis, we establish the consistency of the proposed tests. We also apply the proposed methods to detect changes in mean direction for hourly wind-wave direction (toroidal) measurements and the path (spherical) of the cyclonic storm ``Biporjoy,'' which occurred between 6th and 19th June 2023 over the Arabian Sea, western coast of India.

stat.ME

A Semi-Parametric Torus-to-Torus Regression Model with Geometric Loss: Application to Cyclone Data

This study introduces a novel torus-to-torus regression framework to improve the analysis and prediction of cyclone-driven wind-wave directional dynamics. This research, to our knowledge, establishes a mathematical framework for modeling the regression between bivariate angular predictors and bivariate angular responses for the first time in the literature. The proposed approach enhances the capacity to model coupled directional processes commonly observed in extreme coastal cyclones. The proposed model makes use of generalized Möbius transformation and differential geometry for model building. A new loss function, derived from the intrinsic geometry of the torus, is introduced to facilitate effective semi-parametric estimation without requiring any specific distributional assumptions on the angular error. The prediction error is measured as an angular loss on the surface of the torus and also the angular deflection along normal directions on the unit sphere transported from the torus. Additionally, a new visualization technique for circular data is introduced. The practical relevance of the model is illustrated through its application to wind-wave directional datasets from two major cyclonic events, Amphan and Biparjoy, that impacted the eastern and western coastlines of India, respectively.

stat.ME

Hyperbolic statistical inference for Treatment Effects with Circular biomarker of astigmatism

Circular biomarkers arise naturally in many biomedical applications, particularly in ophthalmology, where angular measurements such as astigmatism are routinely recorded. Similar directional variables also occur in the study of human body rotations, including movements of the hand, waist, neck, and lower limbs. Motivated by a clinical dataset comprising angular measurements of astigmatism induced by two cataract surgery procedures, we propose a novel two-sample testing framework for circular data grounded in hyperbolic geometry. Assuming von Mises distributions with either common or group-specific concentration parameters, we embed the corresponding parameter spaces into the Poincaré disk, an open unit disk endowed with the Poincaré metric.Under this construction, each von Mises distribution is mapped uniquely to a point in the Poincaré disk, yielding a continuous geometric representation that preserves the intrinsic structure of the parameter space. This embedding enables direct comparison of group distributions via hyperbolic distances, leading to natural and interpretable test statistics. We develop permutation-based tests for the common concentration case and bootstrap-based procedures for unequal concentrations. Extensive simulation studies demonstrate stable empirical size, strong consistency, and superior asymptotic power compared with existing competing methods. The proposed methodology is illustrated through a detailed analysis of the cataract surgery dataset, including a clinically informed restructuring of the original observations. The results highlight the practical advantages of incorporating hyperbolic geometry into the analysis of circular biomedical data and underscore the potential of geometry-aware inference for directional biomarkers.

stat.ME

A geometric approach in non-parametric Changepoint detection in circular data

In many temporally ordered data sets, it is observed that the parameters of the underlying distribution change abruptly at unknown times. The detection of such changepoints is important for many applications. While this problem has been studied substantially in the linear data setup, not much work has been done for angular data. In this article, we utilize the intrinsic geometry of a torus to propose new non-parametric tests. First, we propose new tests for the existence of changepoint(s) in the concentration, and second, a test to detect mean direction and/or concentration. The limiting distributions of the test statistics are derived, and their powers are obtained using extensive simulation. It is seen that the tests have better power than the corresponding existing tests. The proposed methods have been implemented on three real-life data sets, revealing interesting insights. In particular, our method, when used to detect simultaneous changes in mean direction and concentration for hourly wind direction measurements of the cyclonic storm "Amphan," identified changepoints that could be associated with important meteorological events.

stat.ME

An Efficient Sampling from Circular Distributions and its Extension to Toroidal Distributions

Sampling from circular distributions is a fundamental task in directional statistics. A key challenge in acceptance-rejection methods lies in selecting an efficient envelope density, as poor choices can lead to low acceptance rates and increased computational cost, especially in large-scale simulations. To address this, we propose a new sampling framework that utilizes the idea of upper Riemann sums to construct a piecewise envelope. This method ensures validity for any Riemann-integrable target density on a bounded interval. This method exhibits enhanced efficacy relative to the present sampling method for the von Mises distribution. Additionally, we introduce a flexible family of distributions defined on the surface of a curved torus, using its area element. The proposed sampling method is then employed to generate samples from the toroidal model. We explore the maximum entropy characterization and other theoretical properties of one of the marginal distributions arising from this construction for the von Mises distribution. To illustrate the practical utility of our framework, we apply the model to a real dataset on wind direction.

stat.ME

Semi-parametric least-area linear-circular regression through Möbius transformation

This paper introduces a novel regression model designed for angular response variables with linear predictors, utilizing a generalized Möbius transformation to define the regression curve. By mapping the real axis to the circle, the model effectively captures the relationship between linear and angular components. A key innovation is the introduction of an area-based loss function, inspired by the geometry of a curved torus, for efficient parameter estimation. The semi-parametric nature of the model eliminates the need for specific distributional assumptions about the angular error, enhancing its versatility. Extensive simulation studies, incorporating von Mises and wrapped Cauchy distributions, highlight the robustness of the framework. The model's practical utility is demonstrated through real-world data analysis of Bitcoin and Ethereum, showcasing its ability to derive meaningful insights from complex data structures.

stat.ME

Intrinsic geometry-inspired dependent toroidal distribution: Application to regression model for astigmatism data

This paper introduces a dependent toroidal distribution, to analyze astigmatism data following cataract surgery. Rather than utilizing the flat torus, we opt to represent the bivariate angular data on the surface of a curved torus, which naturally offers smooth edge identifiability and accommodates a variety of curvatures: positive, negative, and zero. Beginning with the area-uniform toroidal distribution on this curved surface, we develop a five-parameter-dependent toroidal distribution that harnesses its intrinsic geometry via the area element to model the distribution of two dependent circular random variables. We show that both marginal distributions are Cardioid, with one of the conditional variables also following a Cardioid distribution. This key feature enables us to propose a circular-circular regression model based on conditional expectations derived from circular moments. To address the high rejection rate (approximately 50%) in existing acceptance-rejection sampling methods for Cardioid distributions, we introduce an exact sampling method based on a probabilistic transformation. Additionally, we generate random samples from the proposed dependent toroidal distribution through suitable conditioning. This bivariate distribution and the regression model are applied to analyze astigmatism data arising in the follow-up of one and three months due to cataract surgery.

stat.AP

Sampling from the surface of a curved torus: A new genesis

The distributions of toroidal data, often viewed as an extension of circular distributions, do not consider the intrinsic geometry of a curved torus. For the first time, Diaconis et al. (2013)[Diaconis, P., Holmes, S., & Shahshahani, M. (2013). Sampling from a manifold. Advances in modern statistical theory and applications: a Festschrift in honor of Morris L. Eaton, 10, 102-125.] introduce uniform distribution on the surface of a curved torus with respect to its surface area. But the suggested acceptance-rejection method of sampling from it rejects approximately half of the data. We propose a probabilistic transformation for sampling from the same distribution without losing data. In addition, we introduce a new genesis of random samples from some popular circular distributions using histogram-based acceptance-rejection sampling that uses a very thin envelope. The idea leads to generalizing for sampling from distributions on the surface of a curved torus with a high acceptance rate.Apart from reducing computational cost in the inferential study of different toroidal distributions, uniform sampling from the surface of a curve torus will be helpful to understand any unknown distribution on it.

stat.ME

Stochastic ordering results in parallel and series systems with Gumble distributed random variables

The stochastic comparisons of parallel and series system are worthy of study. In this paper, we present some stochastic comparisons of parallel and series systems having independent components from Gumble distribution with two parameters (one location and one shape). Here, we first put a condition for the likelihood ratio ordering of the parallel systems and second we use the concept of vector majorization technique to compare the systems by the reversed hazard rate ordering, the hazard rate ordering, the dispersive ordering, and the less uncertainty ordering with respect to the location parameter.

math.ST

Some ordering properties of highest and lowest order statistics with exponentiated Gumble type-II distributed components

In this paper, we have studied the stochastic comparisons of the highest and lowest order statistics of exponentiated Gumble type-II distribution with three parameters. We have compared both the statistics by using three different stochastic ordering. First, we consider a system with different scale and outer shape parameters and then we study the usual stochastic ordering of the lowest and highest order statistics in the sense of multivariate chain majorization. In addition, we construct two examples to support our results. Second, by using the vector majorization technique, we study the usual stochastic ordering, the reversed failure rate ordering and the likelihood ratio ordering with respect to different outer shape parameters, next, by varying the inner shape parameter, we discuss the usual stochastic order of the lowest order statistics and we have shown that the highest order statistics are not comparable in the usual stochastic ordering by an example.

math.ST

The latent logarithm

Count or non-negative data are often log transformed to improve heteroscedasticity and scaling. To avoid undefined values where the data are zeros, a small pseudocount (e.g. 1) is added across the dataset prior to applying the transformation. This pseudocount considers neither the measured object's a priori abundance nor the confidence with which the measurement was made, making this practice convenient but statistically unfounded. I introduce here the latent logarithm, or lag. lag assumes that each observed measurement is a noisy realization of an unmeasured latent abundance. By taking the logarithm of this learned latent abundance, which reflects both sampling confidence/depth and the object's a priori abundance, lag provides a probabilistically coherent, stable, and intuitive alternative to the questionable, but conventional "log($x$ + pseudocount)."

stat.ME

Learning microbial interaction networks from metagenomic count data

Many microbes associate with higher eukaryotes and impact their vitality. In order to engineer microbiomes for host benefit, we must understand the rules of community assembly and maintenence, which in large part, demands an understanding of the direct interactions between community members. Toward this end, we've developed a Poisson-multivariate normal hierarchical model to learn direct interactions from the count-based output of standard metagenomics sequencing experiments. Our model controls for confounding predictors at the Poisson layer, and captures direct taxon-taxon interactions at the multivariate normal layer using an $\ell_1$ penalized precision matrix. We show in a synthetic experiment that our method handily outperforms state-of-the-art methods such as SparCC and the graphical lasso (glasso). In a real, in planta perturbation experiment of a nine member bacterial community, we show our model, but not SparCC or glasso, correctly resolves a direct interaction structure among three community members that associate with Arabidopsis thaliana roots. We conclude that our method provides a structured, accurate, and distributionally reasonable way of modeling correlated count based random variables and capturing direct interactions among them.

q-bio.QM

Biological Averaging in RNA-Seq

RNA-seq has become a de facto standard for measuring gene expression. Traditionally, RNA-seq experiments are mathematically averaged -- they sequence the mRNA of individuals from different treatment groups, hoping to correlate phenotype with differences in arithmetic read count averages at shared loci of interest. Alternatively, the tissue from the same individuals may be pooled prior to sequencing in what we refer to as a biologically averaged design. As mathematical averaging sequences all individuals it controls for both biological and technical variation; however, is the statistical resolution gained always worth the additional cost? To compare biological and mathematical averaging, we examined theoretical and empirical estimates of statistical efficiency and relative cost efficiency. Though less efficient at a fixed sample size, we found that biological averaging can be more cost efficient than mathematical averaging. With this motivation, we developed a differential expression classifier, ICRBC, that can detect alternatively expressed genes between biologically averaged samples. In simulation studies, we found that biological averaging and subsequent analysis with our classifier performed comparably to existing methods, such as ASC, edgeR, and DESeq, especially when individuals were pooled evenly and less than 20% of the regulome was expected to be differentially regulated. In two technically distinct mouse datasets and one plant dataset, we found that our method was over 87% concordant with edgeR for the 100 most significant features. We therefore conclude biological averaging may sufficiently control biological variation to a level that differences in gene expression may be detectable. In such situations, ICRBC can enable reliable exploratory analysis at a fraction of the cost, especially when interest lies in the most differentially expressed loci.

q-bio.QM