SearcharxivSearch

arXiv subjects

Pritam Ranjan

Publications and source records attributed to Pritam Ranjan.

At least 19 recordsLinked to original sources

smartcor: Intelligent Correlation Method Selection for Mixed Variable Types

Pearson correlation is the default measure of association in most statistical software, yet it is only appropriate for pairs of continuous variables with a linear relationship. When variables are binary, ordinal, or categorical, specialized methods (e.g., point-biserial, polychoric, tetrachoric, and Cram\'{e}r's~$V$) may be more appropriate, but practitioners rarely know which to select. The \textbf{smartcor} package for \textbf{R} (and its companion \textbf{pysmartcor} package for \textbf{Python}) automatically detects variable types, selects the statistically appropriate correlation method for each pair, and explains its reasoning. The package supports 14~correlation and association methods covering all 10~variable-type pair combinations, and distinguishes true correlation (for ordinal and continuous pairs) from statistical association (for nominal categorical pairs). Monte~Carlo simulations validate the selection logic, and a case study with General Social Survey data demonstrates substantive differences between naive and type-aware correlation analysis.

stat.ME

Assessing the influence of social media feedback on traveler's future trip-planning behavior: A multi-model machine learning approach

With the surge of domestic tourism in India and the influence of social media on young tourists, this paper aims to address the research question on how "social return" - responses received on social media sharing - of recent trip details can influence decision-making for short-term future travels. The paper develops a multi-model framework to build a predictive machine learning model that establishes a relationship between a traveler's social return, various social media usage, trip-related factors, and her future trip-planning behavior. The primary data was collected via a survey from Indian tourists. After data cleaning, the imbalance in the data was addressed using a robust oversampling method, and the reliability of the predictive model was ensured by applying a Monte Carlo cross-validation technique. The results suggest at least 75% overall accuracy in predicting the influence of social return on changing the future trip plan. Moreover, the model fit results provide crucial practical implications for the domestic tourism sector in India with future research directions concerning social media, destination marketing, smart tourism, heritage tourism, etc.

cs.SI

Modeling time to failure using a temporal sequence of events

In recent years, the requirement for real-time understanding of machine behavior has become an important objective in industrial sectors to reduce the cost of unscheduled downtime and to maximize production with expected quality. The vast majority of high-end machines are equipped with a number of sensors that can record event logs over time. In this paper, we consider an injection molding (IM) machine that manufactures plastic bottles for soft drink. We have analyzed the machine log data with a sequence of three type of events, ``running with alert'', ``running without alert'', and ``failure''. Failure event leads to downtime of the machine and necessitates maintenance. The sensors are capable of capturing the corresponding operational conditions of the machine as well as the defined states of events. This paper presents a new model to predict a) time to failure of the IM machine and b) identification of important sensors in the system that may explain the events which in-turn leads to failure. The proposed method is more efficient than the popular competitor and can help reduce the downtime costs by controlling operational parameters in advance to prevent failures from occurring too soon.

stat.ME

Solving an Inverse Problem for Time Series Valued Computer Simulators via Multiple Contour Estimation

Computer simulators are often used as a substitute of complex real-life phenomena which are either expensive or infeasible to experiment with. This paper focuses on how to efficiently solve the inverse problem for an expensive to evaluate time series valued computer simulator. The research is motivated by a hydrological simulator which has to be tuned for generating realistic rainfall-runoff measurements in Athens, Georgia, USA. Assuming that the simulator returns g(x,t) over L time points for a given input x, the proposed methodology begins with a careful construction of a discretization (time-) point set (DPS) of size $k << L$, achieved by adopting a regression spline approximation of the target response series at k optimal knots locations $\{t^*_1, t^*_2, ..., t^*_k\}$. Subsequently, we solve k scalar valued inverse problems for simulator $g(x,t^*_j)$ via the contour estimation method. The proposed approach, named MSCE, also facilitates the uncertainty quantification of the inverse solution. Extensive simulation study is used to demonstrate the performance comparison of the proposed method with the popular competitors for several test-function based computer simulators and a real-life rainfall-runoff measurement model.

stat.ME

Valuation of patents in emerging economies: a renewal model-based study of Indian patents

This study uses patent renewal information to estimate private value of patents by technology and ownership status. Patent value refers to the economic reward that the inventor extracts from the patent by making, using or selling an invention. Thus, we measure the value of patent right (private value of patent) from the patentee perspective. Our empirical analysis comprises of 555 patents with application year during 1999 to 2002. The term of these patents either ended in 2018 or lapsed due to non-payment of renewal fee. We model renewal decision of patentee as ordered probit where patent renewal fee increases with the age of patent. Variables such as patent family size, technological scope, number of inventors and grant lag are used as explanatory variables in the corresponding regression. Hence, this paper combines the patentee renewal decision along with patents characteristics and renewal cost schedule to estimate the initial rent distribution. We find that a large number of patents expire at an early stage leaving few patents with high value corroborating the results of studies using European, American and Chinese data. As expected, certain technology class patents enjoy high valuation.

stat.AP

Determinants of Patent Survival in Emerging Economies: Evidence from Residential Patents in India

The purpose of this paper is to use patent level characteristics to estimate the survival of resident patents (filed at the Indian Patent Office (IPO) and assigned to firms in India). This study uses the renewal information of firm-level patents applied during 1st January 1995 and 31st December 2005, which were eventually granted. The data provided by IPO consists of 2025 resident patents assigned to 266 firms (foreign subsidiary firms and domestic firms). The survival analysis is carried out via Kaplan-Meier estimation and Cox proportional hazard regression. The outcomes of this study suggest that the survival length of patents significantly depends on their technological scope and inventor size. Moreover, the patents of the firms taking tax credit benefits exhibit lower survival rate as compared to patents of remaining firms. The study also finds that the patents filed by foreign firms with DSIR affiliation are getting more benefit from the R&D tax incentive policy.

stat.AP

Assessing the Impact of Patent Attributes on the Value of Discrete and Complex Innovations

This study assesses the degree to which the social value of patents can be connected to the private value of patents across discrete and complex innovation. The underlying theory suggests that the social value of cumulative patents is less related to the private value of patents. We use the patents applied between 1995 to 2002 and granted on or before December 2018 from the Indian Patent Office (IPO). Here the patent renewal information is utilized as a proxy for the private value of the patent. We have used a variety of logit regression model for the impact assessment analysis. The results reveal that the technology classification (i.e., discrete versus complex innovations) plays an important role in patent value assessment, and some technologies are significantly different than the others even within the two broader classifications. Moreover, the non-resident patents in India are more likely to have a higher value than the resident patents. According to the conclusions of this study, only a few technologies from the discrete and complex innovation categories have some private value. There is no evidence that patent social value indicators are less useful in complicated technical classes than in discrete ones.

econ.GN

The Evolution of Dynamic Gaussian Process Model with Applications to Malaria Vaccine Coverage Prediction

Gaussian process (GP) based statistical surrogates are popular, inexpensive substitutes for emulating the outputs of expensive computer models that simulate real-world phenomena or complex systems. Here, we discuss the evolution of dynamic GP model - a computationally efficient statistical surrogate for a computer simulator with time series outputs. The main idea is to use a convolution of standard GP models, where the weights are guided by a singular value decomposition (SVD) of the response matrix over the time component. The dynamic GP model also adopts a localized modeling approach for building a statistical model for large datasets. In this chapter, we use several popular test function based computer simulators to illustrate the evolution of dynamic GP models. We also use this model for predicting the coverage of Malaria vaccine worldwide. Malaria is still affecting more than eighty countries concentrated in the tropical belt. In 2019 alone, it was the cause of more than 435,000 deaths worldwide. The malice is easy to cure if diagnosed in time, but the common symptoms make it difficult. We focus on a recently discovered reliable vaccine called Mos-Quirix (RTS,S) which is currently going under human trials. With the help of publicly available data on dosages, efficacy, disease incidence and communicability of other vaccines obtained from the World Health Organisation, we predict vaccine coverage for 78 Malaria-prone countries.

stat.AP

Statistical Modelling and Analysis of the Computer-Simulated Datasets

Over the last two decades, the science has come a long way from relying on only physical experiments and observations to experimentation using computer simulators. This chapter focusses on the modelling and analysis of data arising from computer simulators. It turns out that traditional statistical metamodels are often not very useful for analyzing such datasets. For deterministic computer simulators, the realizations of Gaussian Process (GP) models are commonly used for fitting a surrogate statistical metamodel of the simulator output. The chapter starts with a quick review of the standard GP based statistical surrogate model. The chapter also emphasizes on the numerical instability due to near-singularity of the spatial correlation structure in the GP model fitting process. The authors also present a few generalizations of the GP model, reviews methods and algorithms specifically developed for analyzing big data obtained from computer model runs, and reviews the popular analysis goals of such computer experiments. A few real-life computer simulators are also briefly outlined here.

stat.ME

IsoCheck: An R Package to check Isomorphism for Two-level Factorial Designs with Randomization Restrictions

Factorial designs are often used in various industrial and sociological experiments to identify significant factors and factor combinations that may affect the process response. In the statistics literature, several studies have investigated the analysis, construction, and isomorphism of factorial and fractional factorial designs. When there are multiple choices for a design, it is helpful to have an easy-to-use tool for identifying which are distinct, and which of those can be efficiently analyzed/has good theoretical properties. For this task, we present an R library called IsoCheck that checks the isomorphism of multi-stage 2^n factorial experiments with randomization restrictions. Through representing the factors and their combinations as a finite projective geometry, IsoCheck recasts the problem of searching over all possible relabelings as a search over collineations, then exploits projective geometric properties of the space to make the search much more efficient. Furthermore, a bitstring representation of the factorial effects is used to characterize all possible rearrangements of designs, thus facilitating quick comparisons after relabeling. We present several examples with R code to illustrate the usage of the main functions in IsoCheck. Besides checking equivalence and isomorphism of 2^n multi-stage factorial designs, we demonstrate how the functions of the package can be used to create a catalog of all non-isomorphic designs, and subsequently rank these designs based on a suitably defined ranking criterion. IsoCheck is free software and distributed under the General Public License and available from the Comprehensive R Archive Network.

stat.ME

Inverse Problem for Dynamic Computer Simulators via Multiple Scalar-valued Contour Estimation

In this paper we consider a dynamic computer simulator that produces a time-series response $y_t(x)$ over $L$ time points, for every given input parameter $x$. We propose a method for solving inverse problems, which refer to the finding of a set of inputs that generates a pre-specified simulator output. Inspired by the sequential approach of contour estimation via expected improvement criterion developed by Ranjan et al. (2008, DOI: 10.1198/004017008000000541), our proposed method discretizes the target response series on $k \; (\ll L)$ time points, and then iteratively solves $k$ scalar-valued inverse problems with respect to the discretized targets. We also propose to use spline smoothing of the target response series to identify the optimal number of knots, $k$, and the actual location of the knots for discretization. The performance of the proposed methods is compared for several test-function based computer simulators and the motivating real application that uses a rainfall-runoff measurement model named Matlab-Simulink model.

stat.ME

Isomorphism Check for $2^n$ Factorial Designs with Randomization Restrictions

Factorial designs with randomization restrictions are often used in industrial experiments when a complete randomization of trials is impractical. In the statistics literature, the analysis, construction and isomorphism of factorial designs has been extensively investigated. Much of the work has been on a case-by-case basis -- addressing completely randomized designs, randomized block designs, split-plot designs, etc. separately. In this paper we take a more unified approach, developing theoretical results and an efficient relabeling strategy to both construct and check the isomorphism of multi-stage factorial designs with randomization restrictions. The examples presented in this paper particularly focus on split-lot designs.

stat.ME

A History Matching Approach for Calibrating Hydrological Models

Calibration of hydrological time-series models is a challenging task since these models give a wide spectrum of output series and calibration procedures require significant amount of time. From a statistical standpoint, this model parameter estimation problem simplifies to finding an inverse solution of a computer model that generates pre-specified time-series output (i.e., realistic output series). In this paper, we propose a modified history matching approach for calibrating the time-series rainfall-runoff models with respect to the real data collected from the state of Georgia, USA. We present the methodology and illustrate the application of the algorithm by carrying a simulation study and the two case studies. Several goodness-of-fit statistics were calculated to assess the model performance. The results showed that the proposed history matching algorithm led to a significant improvement, of 30% and 14% (in terms of root mean squared error) and 26% and 118% (in terms of peak percent threshold statistics), for the two case-studies with Matlab-Simulink and SWAT models, respectively.

stat.AP

Global Fitting of the Response Surface via Estimating Multiple Contours of a Simulator

Computer simulators are nowadays widely used to understand complex physical systems in many areas such as aerospace, renewable energy, climate modeling, and manufacturing. One fundamental issue in the study of computer simulators is known as experimental design, that is, how to select the input settings where the computer simulator is run and the corresponding response is collected. Extra care should be taken in the selection process because computer simulators can be computationally expensive to run. The selection shall acknowledge and achieve the goal of the analysis. This article focuses on the goal of producing more accurate prediction which is important for risk assessment and decision making. We propose two new methods of design approaches that sequentially select input settings to achieve this goal. The approaches make novel applications of simultaneous and sequential contour estimations. Numerical examples are employed to demonstrate the effectiveness of the proposed approaches.

stat.ME

A Sequential Design Approach for Calibrating a Dynamic Population Growth Model

A comprehensive understanding of the population growth of a variety of pests is often crucial for efficient crop management. Our motivating application comes from calibrating a two-delay blowfly (TDB) model which is used to simulate the population growth of Panonychus ulmi (Koch) or European red mites that infest on apple leaves and diminish the yield. We focus on the inverse problem, that is, to estimate the set of parameters/inputs of the TDB model that produces the computer model output matching the field observation as closely as possible. The time series nature of both the field observation and the TDB outputs makes the inverse problem significantly more challenging than in the scalar valued simulator case. In spirit, we follow the popular sequential design framework of computer experiments. However, due to the time-series response, a singular value decomposition based Gaussian process model is used for the surrogate model, and subsequently, a new expected improvement criterion is developed for choosing the follow-up points. We also propose a new criterion for extracting the optimal inverse solution from the final surrogate. Three simulated examples and the real-life TDB calibration problem have been used to demonstrate higher accuracy of the proposed approach as compared to popular existing techniques.

stat.ME

Local Gaussian Process Model for Large-scale Dynamic Computer Experiments

The recent accelerated growth in the computing power has generated popularization of experimentation with dynamic computer models in various physical and engineering applications. Despite the extensive statistical research in computer experiments, most of the focus had been on the theoretical and algorithmic innovations for the design and analysis of computer models with scalar responses. In this paper, we propose a computationally efficient statistical emulator for a large-scale dynamic computer simulator (i.e., simulator which gives time series outputs). The main idea is to first find a good local neighbourhood for every input location, and then emulate the simulator output via a singular value decomposition (SVD) based Gaussian process (GP) model. We develop a new design criterion for sequentially finding this local neighbourhood set of training points. Several test functions and a real-life application have been used to demonstrate the performance of the proposed approach over a naive method of choosing local neighbourhood set using the Euclidean distance among design points.

stat.ME

A New Class of Discrete-time Stochastic Volatility Model with Correlated Errors

In an efficient stock market, the returns and their time-dependent volatility are often jointly modeled by stochastic volatility models (SVMs). Over the last few decades several SVMs have been proposed to adequately capture the defining features of the relationship between the return and its volatility. Among one of the earliest SVM, Taylor (1982) proposed a hierarchical model, where the current return is a function of the current latent volatility, which is further modeled as an auto-regressive process. In an attempt to make the SVMs more appropriate for complex realistic market behavior, a leverage parameter was introduced in the Taylor SVM, which however led to the violation of the efficient market hypothesis (EMH, a necessary mean-zero condition for the return distribution that prevents arbitrage possibilities). Subsequently, a host of alternative SVMs had been developed and are currently in use. In this paper, we propose mean-corrections for several generalizations of Taylor SVM that capture the complex market behavior as well as satisfy EMH. We also establish a few theoretical results to characterize the key desirable features of these models, and present comparison with other popular competitors. Furthermore, four real-life examples (Oil price, CITI bank stock price, Euro-USD rate, and S&P 500 index returns) have been used to demonstrate the performance of this new class of SVMs.

stat.AP

Inverse problem for time-series valued computer model via scalarization

For an expensive to evaluate computer simulator, even the estimate of the overall surface can be a challenging problem. In this paper, we focus on the estimation of the inverse solution, i.e., to find the set(s) of input combinations of the simulator that generates (or gives good approximation of) a pre-determined simulator output. Ranjan et al. (2008) proposed an expected improvement criterion under a sequential design framework for the inverse problem with a scalar valued simulator. In this paper, we focus on the inverse problem for a time-series valued simulator. We have used a few simulated and two real examples for performance comparison.

stat.ME