SearcharxivSearch

arXiv subjects

Samiran Ghosh

Publications and source records attributed to Samiran Ghosh.

16 recordsLinked to original sources

Internet malware propagation: Dynamics and control through SEIRV epidemic model with relapse and intervention

Malware attacks in today's vast digital ecosystem pose a serious threat. Understanding malware propagation dynamics and designing effective control strategies are therefore essential. In this work, we propose a generic SEIRV model formulated using ordinary differential equations to study malware spread. We establish the positivity and boundedness of the system, derive the malware propagation threshold, and analyze the local and global stability of the malware-free equilibrium. The separatrix defining epidemic regions in the control space is identified, and the existence of a forward bifurcation is demonstrated. Using normalized forward sensitivity indices, we determine the parameters most influential to the propagation threshold. We further examine the nonlinear dependence of key epidemic characteristics on the transmission rate, including the maximum number of infected, time to peak infection, and total number of infected. We propose a hybrid gradient-based global optimization framework using simulated annealing approach to identify effective and cost-efficient control strategies. Finally, we calibrate the proposed model using infection data from the "Windows Malware Dataset with PE API Calls" and investigated the effect of intervention onset time on averted cases, revealing an exponential decay relationship between delayed intervention and averted cases.

cs.CR

A Longitudinal Measurement Study of Log4Shell Exploitation from a Reactive Network Telescope

The disclosure of the Log4Shell vulnerability in December 2021 led to an unprecedented wave of global scanning and exploitation activity. A recent study provided important initial insights, but was largely limited in duration and geography, focusing primarily on European and U.S. network telescope deployments and covering the immediate aftermath of disclosure. As a result, the longer-term evolution of exploitation behavior and its regional characteristics has remained insufficiently understood. In this paper, we present a longitudinal measurement study of Log4Shell-related traffic observed between December 2021 and October 2025 by a reactive network telescope deployed in India. This vantage point enables examination of sustained exploitation dynamics beyond the initial outbreak phase, including changes in scanning breadth, infrastructure reuse, payload construction, and destination targeting. Our analysis reveals that Log4Shell exploitation persists for several years after disclosure, with activity gradually concentrating around a smaller set of recurring scanner and callback infrastructures, accompanied by an increase in payload obfuscation and shifts in protocol and port usage. A comparative analysis and observations with the benchmark study validate both correlated temporal trends and systematic differences attributable to vantage point placement and coverage. Subsequently, these results demonstrate that Log4Shell remains active well beyond its initial disclosure period, underscoring the value of long-term, geographically diverse measurement for understanding the full lifecycle of critical software vulnerabilities.

cs.CR

Effect of Vaccine Dose Intervals: Considering Immunity Levels, Vaccine Efficacy, and Strain Variants for Disease Control Strategy

In this study, we present an immuno-epidemic model to understand mitigation options during an epidemic break. The model incorporates comorbidity and multiple-vaccine doses through a system of coupled integro-differential equations to analyze the epidemic rate and intensity from a knowledge of the basic reproduction number and time-distributed rate functions. Our modeling results show that the interval between vaccine doses is a key control parameter that can be tuned to significantly influence disease spread. We show that multiple doses induce a hysteresis effect in immunity levels that offers a better mitigation alternative compared to frequent vaccination which is less cost-effective while being more intrusive. Optimal dosing intervals, emphasizing the cost-effectiveness of each vaccination effort, and determined by various factors such as the level of immunity and efficacy of vaccines against different strains, appear to be crucial in disease management. The model is sufficiently generic that can be extended to accommodate specific disease forms.

q-bio.PE

Mathematical Modeling, Analysis and Simulation Utilizing Machine Learning Tools for Assessing the Impact of Climate Lobbying

Climate policy and legislation has a significant influence on both domestic and global responses to the pressing environmental challenges of our time. The effectiveness of such climate legislation is closely tied to the complex dynamics among elected officials, a dynamic significantly shaped by the relentless efforts of lobbying. This project aims to develop a novel compartmental model to forecast the trajectory of climate legislation within the United States. By understanding the dynamics surrounding floor votes, the ramifications of lobbying, and the flow of campaign donations within the chambers of the U.S. Congress, we aim to validate our model through a comprehensive case study of the American Clean Energy and Security Act (ACESA). Our model adeptly captures the nonlinear dynamics among diverse legislative factions, including centrists, ardent supporters, and vocal opponents of the bill, culminating in a rich dynamics of final voting outcomes. We conduct a stability analysis of the model, estimating parameters from public lobbying records and a robust body of existing literature. The numerical verification against the pivotal 2009 ACESA vote, alongside contemporary research, underscores the models promising potential as a tool to understand the dynamics of climate lobbying. We also analyse the pathways of the model that aims to guide future legislative endeavors in the pursuit of effective climate action.

physics.soc-ph

Subgroup analysis in multi level hierarchical cluster randomized trials

Cluster or group randomized trials (CRTs) are increasingly used for both behavioral and system-level interventions, where entire clusters are randomly assigned to a study condition or intervention. Apart from the assigned cluster-level analysis, investigating whether an intervention has a differential effect for specific subgroups remains an important issue, though it is often considered an afterthought in pivotal clinical trials. Determining such subgroup effects in a CRT is a challenging task due to its inherent nested cluster structure. Motivated by a real-life HIV prevention CRT, we consider a three-level cross-sectional CRT, where randomization is carried out at the highest level and subgroups may exist at different levels of the hierarchy. We employ a linear mixed-effects model to estimate the subgroup-specific effects through their maximum likelihood estimators (MLEs). Consequently, we develop a consistent test for the significance of the differential intervention effect between two subgroups at different levels of the hierarchy, which is the key methodological contribution of this work. We also derive explicit formulae for sample size determination to detect a differential intervention effect between two subgroups, aiming to achieve a given statistical power in the case of a planned confirmatory subgroup analysis. The application of our methodology is illustrated through extensive simulation studies using synthetic data, as well as with real-world data from an HIV prevention CRT in The Bahamas.

stat.ME

Estimation of time-varying recovery and death rates from epidemiological data: A new approach

The time-to-recovery or time-to-death for various infectious diseases can vary significantly among individuals, influenced by several factors such as demographic differences, immune strength, medical history, age, pre-existing conditions, and infection severity. To capture these variations, time-since-infection dependent recovery and death rates offer a detailed description of the epidemic. However, obtaining individual-level data to estimate these rates is challenging, while aggregate epidemiological data (such as the number of new infections, number of active cases, number of new recoveries, and number of new deaths) are more readily available. In this article, a new methodology is proposed to estimate time-since-infection dependent recovery and death rates using easily available data sources, accommodating irregular data collection timings reflective of real-world reporting practices. The Nadaraya-Watson estimator is utilized to derive the number of new infections. This model improves the accuracy of epidemic progression descriptions and provides clear insights into recovery and death distributions. The proposed methodology is validated using COVID-19 data and its general applicability is demonstrated by applying it to some other diseases like measles and typhoid.

stat.AP

An age-distributed immuno-epidemiological model with information-based vaccination decision

A new age-distributed immuno-epidemiological model with information-based vaccine uptake suggested in this work represents a system of integro-differential equations for the numbers of susceptible individuals, infected individuals, vaccinated individuals and recovered individuals. This model describes the influence of vaccination decision on epidemic progression in different age groups. We prove the existence and uniqueness of a positive solution using the fixed point theory. In a particular case of age-independent model, we determine the final size of epidemic, that is, the limiting number of susceptible individuals at asymptotically large time. Numerical simulations show that the information-based vaccine acceptance can significantly influence the epidemic progression. Though the initial stage of epidemic progression is the same for all memory kernels, as the epidemic progresses and more information about the disease becomes available, further epidemic progression strongly depends on the memory effect. Short-range memory kernel appears to be more effective in restraining the epidemic outbreaks because it allows for more responsive and adaptive vaccination decisions based on the most recent information about the disease.

q-bio.PE

Spatio-temporal chaos and clustering induced by nonlocal information and vaccine hesitancy in the SIR epidemic model

Human behavior, and in particular vaccine hesitancy, is a critical factor for the control of childhood infectious disease. Here we propose a spatio-temporal behavioral epidemiology model where the vaccine propensity depends on information that is non-local in space and in time. The properties of the proposed model are analysed under different hypotheses on the spatio-temporal kernels tuning the vaccination response of individuals. As a main result, we could numerically show that vaccine hesitancy induces the onset of many dynamic patterns of relevance for epidemiology. In particular we observed: behavior-modulated patterns and spatio-temporal chaos. This is the first known example of human behavior-induced spatio-temporal chaos in statistical physics of vaccination. Patterns and spatio-temporal chaos are difficult to deal with, from the Public Health viewpoint, hence showing that vaccine hesitancy can cause them could be of interest. Additionally, we propose a new simple heuristic algorithm to estimate the Maximum Lyapunov Exponent.

physics.soc-ph

Survival trees for right-censored data based on score based parameter instability test

Survival analysis of right censored data arises often in many areas of research including medical research. Effect of covariates (and their interactions) on survival distribution can be studied through existing methods which requires to pre-specify the functional form of the covariates including their interactions. Survival trees offer relatively flexible approach when the form of covariates' effects is unknown. Most of the currently available survival tree construction techniques are not based on a formal test of significance; however, recently proposed ctree algorithm (Hothorn et al., 2006) uses permutation test for splitting decision that may be conservative at times. We consider parameter instability test of statistical significance of heterogeneity to guard against spurious findings of variation in covariates' effect without being overly conservative. We have proposed SurvCART algorithm to construct survival tree under conditional inference framework (Hothorn et al., 2006) that selects splitting variable via parameter instability test and subsequently finds the optimal split based on some maximally chosen statistic. Notably, unlike the existing algorithms which focuses only on heterogeneity in event time distribution, the proposed SurvCART algorithm can take splitting decision based in censoring distribution as well along with heterogeneity in event time distribution. The operating characteristics of parameter instability test and comparative assessment of SurvCART algorithm were carried out via simulation. Finally, SurvCART algorithm was applied to a real data setting. The proposed method is fully implemented in R package LongCART available on CRAN.

stat.ME

Robust Variable Selection Criteria for the Penalized Regression

We propose a robust variable selection procedure using a divergence based M-estimator combined with a penalty function. It produces robust estimates of the regression parameters and simultaneously selects the important explanatory variables. An efficient algorithm based on the quadratic approximation of the estimating equation is constructed. The asymptotic distribution and the influence function of the regression coefficients are derived. The widely used model selection procedures based on the Mallows's $C_p$ statistic and Akaike information criterion (AIC) often show very poor performance in the presence of heavy-tailed error or outliers. For this purpose, we introduce robust versions of these information criteria based on our proposed method. The simulation studies show that the robust variable selection technique outperforms the classical likelihood-based techniques in the presence of outliers. The performance of the proposed method is also explored through the real data analysis.

stat.ME

Shock waves in a rotating non-Maxwellian magnetized dusty plasma

A theoretical model is presented to study characteristics of dust acoustic shock in a viscous, magnetized and rotating dusty plasma at both fast and slow time scales. By employing reductive perturbation technique the nonlinear Zakharov--Kuznetsov (ZK) equation has been derived for both cases when dust is inactive and dynamic (fast and slow time scales). Both electrons and ions are considered to follow kappa/Cairns distribution. It is observed that the viscosity in both cases when dust is in background and active plays as a key role in dissipation for the propagation of acoustic shock. Magnetic field and rotation are responsible for the dispersive term. Superthermality has been found to affect significantly on the formation of shock wave along with viscous nature of plasma. The present investigation may be beneficial to understanding the rotating plasma in particular experiments being carried out.

physics.plasm-ph

Qualitative analysis of certain generalized classes of quadratic oscillator systems

We carry out a systematic qualitative analysis of the two quadratic schemes of generalized oscillators recently proposed by C. Quesne [J.Math.Phys.\textbf{56},012903 (2015)]. By performing a local analysis of the governing potentials we demonstrate that while the first potential admits a pair of equilibrium points one of which is typically a center for both signs of the coupling strength $λ$, the other points to a centre for $λ< 0$ but a saddle $λ> 0$. On the other hand, the second potential reveals only a center for both the signs of $λ$ from a linear stability analysis. We carry out our study by extending Quesne's scheme to include the effects of a linear dissipative term. An important outcome is that we run into a remarkable transition to chaos in the presence of a periodic force term $f\cos ωt$.

math-ph

Nonlinear Dynamics of a position-dependent mass driven Duffing-type oscillator

We examine some nontrivial consequences that emerge from interpreting a position-dependent mass (PDM) driven Duffing oscillator in the presence of a quartic potential. The propagation dynamics is studied numerically and sensi- tivity to the PDM-index is noted. Remarkable transitions from a limit cycle to chaos through period doubling and from a chaotic to a regular motion through intermediate periodic and chaotic routes are demonstrated.

math-ph

Outlier Detection Techniques for SQL and ETL Tuning

RDBMS is the heart for both OLTP and OLAP types of applications. For both types of applications thousands of queries expressed in terms of SQL are executed on daily basis. All the commercial DBMS engines capture various attributes in system tables about these executed queries. These queries need to conform to best practices and need to be tuned to ensure optimal performance. While we use checklists, often tools to enforce the same, a black box technique on the queries for profiling, outlier detection is not employed for a summary level understanding. This is the motivation of the paper, as this not only points out to inefficiencies built in the system, but also has the potential to point evolving best practices and inappropriate usage. Certainly this can reduce latency in information flow and optimal utilization of hardware and software capacity. In this paper we start with formulating the problem. We explore four outlier detection techniques. We apply these techniques over rich corpora of production queries and analyze the results. We also explore benefit of an ensemble approach. We conclude with future courses of action. The same philosophy we have used for optimization of extraction, transform, load (ETL) jobs in one of our previous work. We give a brief introduction of the same in section four.

cs.DB

Outlier detection from ETL Execution trace

Extract, Transform, Load (ETL) is an integral part of Data Warehousing (DW) implementation. The commercial tools that are used for this purpose captures lot of execution trace in form of various log files with plethora of information. However there has been hardly any initiative where any proactive analyses have been done on the ETL logs to improve their efficiency. In this paper we utilize outlier detection technique to find the processes varying most from the group in terms of execution trace. As our experiment was carried on actual production processes, any outlier we would consider as a signal rather than a noise. To identify the input parameters for the outlier detection algorithm we employ a survey among developer community with varied mix of experience and expertise. We use simple text parsing to extract these features from the logs, as shortlisted from the survey. Subsequently we applied outlier detection technique (Clustering based) on the logs. By this process we reduced our domain of detailed analysis from 500 logs to 44 logs (8 Percentage). Among the 5 outlier cluster, 2 of them are genuine concern, while the other 3 figure out because of the huge number of rows involved.

cs.DB

An imputation-based approach for parameter estimation in the presence of ambiguous censoring with application in industrial supply chain

This paper describes a novel approach based on "proportional imputation" when identical units produced in a batch have random but independent installation and failure times. The current problem is motivated by a real life industrial production-delivery supply chain where identical units are shipped after production to a third party warehouse and then sold at a future date for possible installation. Due to practical limitations, at any given time point, the exact installation as well as the failure times are known for only those units which have failed within that time frame after the installation. Hence, in-house reliability engineers are presented with a very limited, as well as partial, data to estimate different model parameters related to installation and failure distributions. In reality, other units in the batch are generally not utilized due to lack of proper statistical methodology, leading to gross misspecification. In this paper we have introduced a likelihood based parametric and computationally efficient solution to overcome this problem.

stat.AP