Searcharxiv⌕ Search

arXiv subjects

Sandeep Juneja

Publications and source records attributed to Sandeep Juneja.

35 records · Page 2Linked to original sources

Modelling the Second Covid-19 Wave in Mumbai

India has been hit by a huge second wave of Covid-19 that started in mid-February 2021. Mumbai was amongst the first cities to see the increase. In this report, we use our agent based simulator to computationally study the second wave in Mumbai. We build upon our earlier analysis, where projections were made from November 2020 onwards. We use our simulator to conduct an extensive scenario analysis - we play out many plausible scenarios through varying economic activity, reinfection levels, population compliance, infectiveness, prevalence and lethality of the possible variant strains, and infection spread via local trains to arrive at those that may better explain the second wave fatality numbers. We observe and highlight that timings of peak and valley of the fatalities in the second wave are robust to many plausible scenarios, suggesting that they are likely to be accurate projections for Mumbai. During the second wave, the observed fatalities were low in February and mid-March and saw a phase change or a steep increase in the growth rate after around late March. We conduct extensive experiments to replicate this observed sharp convexity. This is not an easy phenomena to replicate, and we find that explanations such as increased laxity in the population, increased reinfections, increased intensity of infections in Mumbai transportation, increased lethality in the virus, or a combination amongst them, generally do a poor job of matching this pattern. We find that the most likely explanation is presence of small amount of extremely infective variant on February 1 that grows rapidly thereafter and becomes a dominant strain by Mid-March. From a prescriptive view, this points to an urgent need for extensive and continuous genome sequencing to establish existence and prevalence of different virus strains in Mumbai and in India, as they evolve over time.

q-bio.PE↗

Regret Minimization in Heavy-Tailed Bandits

We revisit the classic regret-minimization problem in the stochastic multi-armed bandit setting when the arm-distributions are allowed to be heavy-tailed. Regret minimization has been well studied in simpler settings of either bounded support reward distributions or distributions that belong to a single parameter exponential family. We work under the much weaker assumption that the moments of order $(1+ε)$ are uniformly bounded by a known constant B, for some given $ε> 0$. We propose an optimal algorithm that matches the lower bound exactly in the first-order term. We also give a finite-time bound on its regret. We show that our index concentrates faster than the well known truncated or trimmed empirical mean estimators for the mean of heavy-tailed distributions. Computing our index can be computationally demanding. To address this, we develop a batch-based algorithm that is optimal up to a multiplicative constant depending on the batch size. We hence provide a controlled trade-off between statistical optimality and computational cost.

cs.LG↗

COVID-19 Epidemic in Mumbai: Projections, full economic opening, and containment zones versus contact tracing and testing: An Update

Mumbai, amongst the most densely populated cities in the world, has witnessed the fourth largest number of cases and the largest number of deaths among all the cities in India (as of 28th October 2020). Along with the rest of India, lockdowns (of varying degrees) have been in effect in Mumbai since March 25, 2020. Given the large economic toll on the country from the lockdown and the related restrictions on mobility of people and goods, swift opening of the economy especially in a financial hub such as Mumbai becomes critical. In this report, we use the IISc-TIFR agent based simulator to develop long term projections for Mumbai under realistic scenarios related to Mumbai's opening of the workplaces, or equivalently, the economy, and the associated public transportation through local trains and buses. These projections were developed taking into account a possible second wave if the economy and the local trains are fully opened either on November 1, 2020 or on January 1, 2021. The impact on infection spread in Mumbai if the schools and colleges open on January first week 2021 is also considered. We also try to account for the increased intermingling amongst the population during the Ganeshotsav festival as well as around the Navratri/Dussehra and Diwali festival. Our conclusion, based on our simulations, is that the impact of fully opening up the economy on November 1 is manageable provided reasonable medical infrastructure is in place. Further, schools and colleges opening in January do not lead to excessive increase in infections. The report also explores the relative effectiveness of contact tracing vs containment zones, and also includes very rudimentary results of the effect of vaccinating the elderly population in February 2021.

physics.soc-ph↗

City-Scale Agent-Based Simulators for the Study of Non-Pharmaceutical Interventions in the Context of the COVID-19 Epidemic

We highlight the usefulness of city-scale agent-based simulators in studying various non-pharmaceutical interventions to manage an evolving pandemic. We ground our studies in the context of the COVID-19 pandemic and demonstrate the power of the simulator via several exploratory case studies in two metropolises, Bengaluru and Mumbai. Such tools become common-place in any city administration's tool kit in our march towards digital health.

q-bio.PE↗

COVID-19 Epidemic Study II: Phased Emergence From the Lockdown in Mumbai

The nation-wide lockdown starting 25 March 2020, aimed at suppressing the spread of the COVID-19 disease, was extended until 31 May 2020 in three subsequent orders by the Government of India. The extended lockdown has had significant social and economic consequences and `lockdown fatigue' has likely set in. Phased reopening began from 01 June 2020 onwards. Mumbai, one of the most crowded cities in the world, has witnessed both the largest number of cases and deaths among all the cities in India (41986 positive cases and 1368 deaths as of 02 June 2020). Many tough decisions are going to be made on re-opening in the next few days. In an earlier IISc-TIFR Report, we presented an agent-based city-scale simulator(ABCS) to model the progression and spread of the infection in large metropolises like Mumbai and Bengaluru. As discussed in IISc-TIFR Report 1, ABCS is a useful tool to model interactions of city residents at an individual level and to capture the impact of non-pharmaceutical interventions on the infection spread. In this report we focus on Mumbai. Using our simulator, we consider some plausible scenarios for phased emergence of Mumbai from the lockdown, 01 June 2020 onwards. These include phased and gradual opening of the industry, partial opening of public transportation (modelling of infection spread in suburban trains), impact of containment zones on controlling infections, and the role of compliance with respect to various intervention measures including use of masks, case isolation, home quarantine, etc. The main takeaway of our simulation results is that a phased opening of workplaces, say at a conservative attendance level of 20 to 33\%, is a good way to restart economic activity while ensuring that the city's medical care capacity remains adequate to handle the possible rise in the number of COVID-19 patients in June and July.

q-bio.PE↗

Discriminative Learning via Adaptive Questioning

We consider the problem of designing an adaptive sequence of questions that optimally classify a candidate's ability into one of several categories or discriminative grades. A candidate's ability is modeled as an unknown parameter, which, together with the difficulty of the question asked, determines the likelihood with which s/he is able to answer a question correctly. The learning algorithm is only able to observe these noisy responses to its queries. We consider this problem from a fixed confidence-based $δ$-correct framework, that in our setting seeks to arrive at the correct ability discrimination at the fastest possible rate while guaranteeing that the probability of error is less than a pre-specified and small $δ$. In this setting we develop lower bounds on any sequential questioning strategy and develop geometrical insights into the problem structure both from primal and dual formulation. In addition, we arrive at algorithms that essentially match these lower bounds. Our key conclusions are that, asymptotically, any candidate needs to be asked questions at most at two (candidate ability-specific) levels, although, in a reasonably general framework, questions need to be asked only at a single level. Further, and interestingly, the problem structure facilitates endogenous exploration, so there is no need for a separately designed exploration stage in the algorithm.

cs.LG↗

Credit Risk: Simple Closed Form Approximate Maximum Likelihood Estimator

We consider discrete default intensity based and logit type reduced form models for conditional default probabilities for corporate loans where we develop simple closed form approximations to the maximum likelihood estimator (MLE) when the underlying covariates follow a stationary Gaussian process. In a practically reasonable asymptotic regime where the default probabilities are small, say 1-3% annually, the number of firms and the time period of data available is reasonably large, we rigorously show that the proposed estimator behaves similarly or slightly worse than the MLE when the underlying model is correctly specified. For more realistic case of model misspecification, both estimators are seen to be equally good, or equally bad. Further, beyond a point, both are more-or-less insensitive to increase in data. These conclusions are validated on empirical and simulated data. The proposed approximations should also have applications outside finance, where logit-type models are used and probabilities of interest are small.

econ.EM↗

Unbiased Estimation of the Reciprocal Mean for Non-negative Random Variables

Many simulation problems require the estimation of a ratio of two expectations. In recent years Monte Carlo estimators have been proposed that can estimate such ratios without bias. We investigate the theoretical properties of such estimators for the estimation of $β= 1/\mathbb{E}\, Z$, where $Z \geq 0$. The estimator, $\widehat β(w)$, is of the form $w/f_w(N) \prod_{i=1}^N (1 - w\, Z_i)$, where $w < 2β$ and $N$ is any random variable with probability mass function $f_w$ on the positive integers. For a fixed $w$, the optimal choice for $f_w$ is well understood, but less so the choice of $w$. We study the properties of $\widehat β(w)$ as a function of~$w$ and show that its expected time variance product decreases as $w$ decreases, even though the cost of constructing the estimator increases with $w$. We also show that the estimator is asymptotically equivalent to the maximum likelihood (biased) ratio estimator and establish practical confidence intervals.

math.ST↗

Sample complexity of partition identification using multi-armed bandits

Given a vector of probability distributions, or arms, each of which can be sampled independently, we consider the problem of identifying the partition to which this vector belongs from a finitely partitioned universe of such vector of distributions. We study this as a pure exploration problem in multi armed bandit settings and develop sample complexity bounds on the total mean number of samples required for identifying the correct partition with high probability. This framework subsumes well studied problems such as finding the best arm or the best few arms. We consider distributions belonging to the single parameter exponential family and primarily consider partitions where the vector of means of arms lie either in a given set or its complement. The sets considered correspond to distributions where there exists a mean above a specified threshold, where the set is a half space and where either the set or its complement is a polytope, or more generally, a convex set. In these settings, we characterize the lower bounds on mean number of samples for each arm highlighting their dependence on the problem geometry. Further, inspired by the lower bounds, we propose algorithms that can match these bounds asymptotically with decreasing probability of error. Applications of this framework may be diverse. We briefly discuss one associated with finance.

cs.LG↗

Selecting the best system and multi-armed bandits

Consider the problem of finding a population or a probability distribution amongst many with the largest mean when these means are unknown but population samples can be simulated or otherwise generated. Typically, by selecting largest sample mean population, it can be shown that false selection probability decays at an exponential rate. Lately, researchers have sought algorithms that guarantee that this probability is restricted to a small $δ$ in order $\log(1/δ)$ computational time by estimating the associated large deviations rate function via simulation. We show that such guarantees are misleading when populations have unbounded support even when these may be light-tailed. Specifically, we show that any policy that identifies the correct population with probability at least $1-δ$ for each problem instance requires infinite number of samples in expectation in making such a determination in any problem instance. This suggests that some restrictions are essential on populations to devise $O(\log(1/δ))$ algorithms with $1 - δ$ correctness guarantees. We note that under restriction on population moments, such methods are easily designed, and that sequential methods from stochastic multi-armed bandit literature can be adapted to devise such algorithms.

math.PR↗

Path-ZVA: general, efficient and automated importance sampling for highly reliable Markovian systems

We introduce Path-ZVA: an efficient simulation technique for estimating the probability of reaching a rare goal state before a regeneration state in a (discrete-time) Markov chain. Standard Monte Carlo simulation techniques do not work well for rare events, so we use importance sampling; i.e., we change the probability measure governing the Markov chain such that transitions `towards' the goal state become more likely. To do this we need an idea of distance to the goal state, so some level of knowledge of the Markov chain is required. In this paper, we use graph analysis to obtain this knowledge. In particular, we focus on knowledge of the shortest paths (in terms of `rare' transitions) to the goal state. We show that only a subset of the (possibly huge) state space needs to be considered. This is effective when the high dependability of the system is primarily due to high component reliability, but less so when it is due to high redundancies. For several models we compare our results to well-known importance sampling methods from the literature and demonstrate the large potential gains of our method.

math.PR↗

Exact and efficient simulation of tail probabilities of heavy-tailed infinite series

We develop an efficient simulation algorithm for computing the tail probabilities of the infinite series $S = \sum_{n \geq 1} a_n X_n$ when random variables $X_n$ are heavy-tailed. As $S$ is the sum of infinitely many random variables, any simulation algorithm that stops after simulating only fixed, finitely many random variables is likely to introduce a bias. We overcome this challenge by rewriting the tail probability of interest as a sum of a random number of telescoping terms, and subsequently developing conditional Monte Carlo based low variance simulation estimators for each telescoping term. The resulting algorithm is proved to result in estimators that a) have no bias, and b) require only a fixed, finite number of replications irrespective of how rare the tail probability of interest is. Thus, by combining a traditional variance reduction technique such as conditional Monte Carlo with more recent use of auxiliary randomization to remove bias in a multi-level type representation, we develop an efficient and unbiased simulation algorithm for tail probabilities of $S$. These have many applications including in analysis of financial time-series and stochastic recurrence equations arising in models in actuarial risk and population biology.

math.PR↗

Incorporating Views on Marginal Distributions in the Calibration of Risk Models

Entropy based ideas find wide-ranging applications in finance for calibrating models of portfolio risk as well as options pricing. The abstracted problem, extensively studied in the literature, corresponds to finding a probability measure that minimizes relative entropy with respect to a specified measure while satisfying constraints on moments of associated random variables. These moments may correspond to views held by experts in the portfolio risk setting and to market prices of liquid options for options pricing models. However, it is reasonable that in the former settings, the experts may have views on tails of risks of some securities. Similarly, in options pricing, significant literature focuses on arriving at the implied risk neutral density of benchmark instruments through observed market prices. With the intent of calibrating models to these more general stipulations, we develop a unified entropy based methodology to allow constraints on both moments as well as marginal distributions of functions of underlying securities. This is applied to Markowitz portfolio framework, where a view that a particular portfolio incurs heavy tailed losses is shown to lead to fatter and more reasonable tails for losses of component securities. We also use this methodology to price non-traded options using market information such as observed option prices and implied risk neutral densities of benchmark instruments.

q-fin.ST↗

State-independent Importance Sampling for Random Walks with Regularly Varying Increments

We develop importance sampling based efficient simulation techniques for three commonly encountered rare event probabilities associated with random walks having i.i.d. regularly varying increments; namely, 1) the large deviation probabilities, 2) the level crossing probabilities, and 3) the level crossing probabilities within a regenerative cycle. Exponential twisting based state-independent methods, which are effective in efficiently estimating these probabilities for light-tailed increments are not applicable when the increments are heavy-tailed. To address the latter case, more complex and elegant state-dependent efficient simulation algorithms have been developed in the literature over the last few years. We propose that by suitably decomposing these rare event probabilities into a dominant and further residual components, simpler state-independent importance sampling algorithms can be devised for each component resulting in composite unbiased estimators with desirable efficiency properties. When the increments have infinite variance, there is an added complexity in estimating the level crossing probabilities as even the well known zero-variance measures have an infinite expected termination time. We adapt our algorithms so that this expectation is finite while the estimators remain strongly efficient. Numerically, the proposed estimators perform at least as well, and sometimes substantially better than the existing state-dependent estimators in the literature.

math.PR↗

Regenerative Simulation for Queueing Networks with Exponential or Heavier Tail Arrival Distributions

Multiclass open queueing networks find wide applications in communication, computer and fabrication networks. Often one is interested in steady-state performance measures associated with these networks. Conceptually, under mild conditions, a regenerative structure exists in multiclass networks, making them amenable to regenerative simulation for estimating the steady-state performance measures. However, typically, identification of a regenerative structure in these networks is difficult. A well known exception is when all the interarrival times are exponentially distributed, where the instants corresponding to customer arrivals to an empty network constitute a regenerative structure. In this paper, we consider networks where the interarrival times are generally distributed but have exponential or heavier tails. We show that these distributions can be decomposed into a mixture of sums of independent random variables such that at least one of the components is exponentially distributed. This allows an easily implementable embedded regenerative structure in the Markov process. We show that under mild conditions on the network primitives, the regenerative mean and standard deviation estimators are consistent and satisfy a joint central limit theorem useful for constructing asymptotically valid confidence intervals. We also show that amongst all such interarrival time decompositions, the one with the largest mean exponential component minimizes the asymptotic variance of the standard deviation estimator.

math.PR↗

Efficient simulation of density and probability of large deviations of sum of random vectors using saddle point representations

We consider the problem of efficient simulation estimation of the density function at the tails, and the probability of large deviations for a sum of independent, identically distributed, light-tailed and non-lattice random vectors. The latter problem besides being of independent interest, also forms a building block for more complex rare event problems that arise, for instance, in queueing and financial credit risk modelling. It has been extensively studied in literature where state independent exponential twisting based importance sampling has been shown to be asymptotically efficient and a more nuanced state dependent exponential twisting has been shown to have a stronger bounded relative error property. We exploit the saddle-point based representations that exist for these rare quantities, which rely on inverting the characteristic functions of the underlying random vectors. These representations reduce the rare event estimation problem to evaluating certain integrals, which may via importance sampling be represented as expectations. Further, it is easy to identify and approximate the zero-variance importance sampling distribution to estimate these integrals. We identify such importance sampling measures and show that they possess the asymptotically vanishing relative error property that is stronger than the bounded relative error property. To illustrate the broader applicability of the proposed methodology, we extend it to similarly efficiently estimate the practically important expected overshoot of sums of iid random variables.

math.PR↗

Incorporating fat tails in financial models using entropic divergence measures

In the existing financial literature, entropy based ideas have been proposed in portfolio optimization, in model calibration for options pricing as well as in ascertaining a pricing measure in incomplete markets. The abstracted problem corresponds to finding a probability measure that minimizes the relative entropy (also called $I$-divergence) with respect to a known measure while it satisfies certain moment constraints on functions of underlying assets. In this paper, we show that under $I$-divergence, the optimal solution may not exist when the underlying assets have fat tailed distributions, ubiquitous in financial practice. We note that this drawback may be corrected if `polynomial-divergence' is used. This divergence can be seen to be equivalent to the well known (relative) Tsallis or (relative) Renyi entropy. We discuss existence and uniqueness issues related to this new optimization problem as well as the nature of the optimal solution under different objectives. We also identify the optimal solution structure under $I$-divergence as well as polynomial-divergence when the associated constraints include those on marginal distribution of functions of underlying assets. These results are applied to a simple problem of model calibration to options prices as well as to portfolio modeling in Markowitz framework, where we note that a reasonable view that a particular portfolio of assets has heavy tailed losses may lead to fatter and more reasonable tail distributions of all assets.

q-fin.ST↗