SearcharxivSearch

arXiv subjects

M. Ashok Kumar

Publications and source records attributed to M. Ashok Kumar.

16 recordsLinked to original sources

A Generalized Cram\'er-Rao Bound Using Information Geometry

In information geometry, statistical models are considered as differentiable manifolds, where each probability distribution represents a unique point on the manifold. A Riemannian metric can be systematically obtained from a divergence function using Eguchi's theory (1992); the well-known Fisher-Rao metric is obtained from the Kullback-Leibler (KL) divergence. The geometric derivation of the classical Cram\'er-Rao Lower Bound (CRLB) by Amari and Nagaoka (2000) is based on this metric. In this paper, we study a Riemannian metric obtained by applying Eguchi's theory to the Basu-Harris-Hjort-Jones (BHHJ) divergence (1998) and derive a generalized Cram\'er-Rao bound using Amari-Nagaoka's approach. There are potential applications for this bound in robust estimation.

math.ST

Generalized Fisher-Darmois-Koopman-Pitman Theorem and Rao-Blackwell Type Estimators for Power-Law Distributions

This paper generalizes the notion of sufficiency for estimation problems beyond maximum likelihood. In particular, we consider estimation problems based on Jones et al. and Basu et al. likelihood functions that are popular among distance-based robust inference methods. We first characterize the probability distributions that always have a fixed number of sufficient statistics (independent of sample size) with respect to these likelihood functions. These distributions are power-law extensions of the usual exponential family and contain Student distributions as a special case. We then extend the notion of minimal sufficient statistics and compute it for these power-law families. Finally, we establish a Rao-Blackwell-type theorem for finding the best estimators for a power-law family. This helps us establish Cramér-Rao-type lower bounds for power-law families.

math.ST

Information Geometry for the Working Information Theorist

Information geometry is a study of statistical manifolds, that is, spaces of probability distributions from a geometric perspective. Its classical information-theoretic applications relate to statistical concepts such as Fisher information, sufficient statistics, and efficient estimators. Today, information geometry has emerged as an interdisciplinary field that finds applications in diverse areas such as radar sensing, array signal processing, quantum physics, deep learning, and optimal transport. This article presents an overview of essential information geometry to initiate an information theorist, who may be unfamiliar with this exciting area of research. We explain the concepts of divergences on statistical manifolds, generalized notions of distances, orthogonality, and geodesics, thereby paving the way for concrete applications and novel theoretical investigations. We also highlight some recent information-geometric developments, which are of interest to the broader information theory community.

cs.IT

Information Geometry and Classical Cramér-Rao Type Inequalities

We examine the role of information geometry in the context of classical Cramér-Rao (CR) type inequalities. In particular, we focus on Eguchi's theory of obtaining dualistic geometric structures from a divergence function and then applying Amari-Nagoaka's theory to obtain a CR type inequality. The classical deterministic CR inequality is derived from Kullback-Leibler (KL)-divergence. We show that this framework could be generalized to other CR type inequalities through four examples: $α$-version of CR inequality, generalized CR inequality, Bayesian CR inequality, and Bayesian $α$-CR inequality. These are obtained from, respectively, $I_α$-divergence (or relative $α$-entropy), generalized Csiszár divergence, Bayesian KL divergence, and Bayesian $I_α$-divergence.

cs.IT

Projection Theorems and Estimating Equations for Power-Law Models

We extend projection theorems concerning Hellinger and Jones et al. divergences to the continuous case. These projection theorems reduce certain estimation problems on generalized exponential models to linear problems. We introduce the notion of regularity for generalized exponential models and show that the projection theorems in this case are similar to the ones in discrete and canonical case. We also apply these ideas to solve certain estimation problems concerning Student and Cauchy distributions.

math.ST

Are Guessing, Source Coding, and Tasks Partitioning Birds of a Feather?

This paper establishes a close relationship among the four information theoretic problems, namely Campbell source coding, Arikan guessing, Huleihel et al. memoryless guessing and Bunte and Lapidoth tasks partitioning problems. We first show that the aforementioned problems are mathematically related via a general moment minimization problem whose optimum solution is given in terms of Renyi entropy. We then propose a general framework for the mismatched version of these problems and establish all the asymptotic results using this framework. Further, we study an ordered tasks partitioning problem that turns out to be a generalisation of Arikan's guessing problem. Finally, with the help of this general framework, we establish an equivalence among all these problems, in the sense that, knowing an asymptotically optimal solution in one problem helps us find the same in all other problems.

cs.IT

Cramér-Rao Lower Bounds Arising from Generalized Csiszár Divergences

We study the geometry of probability distributions with respect to a generalized family of Csiszár $f$-divergences. A member of this family is the relative $α$-entropy which is also a Rényi analog of relative entropy in information theory and known as logarithmic or projective power divergence in statistics. We apply Eguchi's theory to derive the Fisher information metric and the dual affine connections arising from these generalized divergence functions. This enables us to arrive at a more widely applicable version of the Cramér-Rao inequality, which provides a lower bound for the variance of an estimator for an escort of the underlying parametric probability distribution. We then extend the Amari-Nagaoka's dually flat structure of the exponential and mixer models to other distributions with respect to the aforementioned generalized metric. We show that these formulations lead us to find unbiased and efficient estimators for the escort model. Finally, we compare our work with prior results on generalized Cramér-Rao inequalities that were derived from non-information-geometric frameworks.

cs.IT

Generalized Bayesian Cramér-Rao Inequality via Information Geometry of Relative $α$-Entropy

The relative $α$-entropy is the Rényi analog of relative entropy and arises prominently in information-theoretic problems. Recent information geometric investigations on this quantity have enabled the generalization of the Cramér-Rao inequality, which provides a lower bound for the variance of an estimator of an escort of the underlying parametric probability distribution. However, this framework remains unexamined in the Bayesian framework. In this paper, we propose a general Riemannian metric based on relative $α$-entropy to obtain a generalized Bayesian Cramér-Rao inequality. This establishes a lower bound for the variance of an unbiased estimator for the $α$-escort distribution starting from an unbiased estimator for the underlying distribution. We show that in the limiting case when the entropy order approaches unity, this framework reduces to the conventional Bayesian Cramér-Rao inequality. Further, in the absence of priors, the same framework yields the deterministic Cramér-Rao inequality.

cs.IT

A Unified Framework for Problems on Guessing, Source Coding and Task Partitioning

We study four problems namely, Campbell's source coding problem, Arikan's guessing problem, Huieihel et al.'s memoryless guessing problem, and Bunte and Lapidoth's task partitioning problem. We observe a close relationship among these problems. In all these problems, the objective is to minimize moments of some functions of random variables, and Rényi entropy and Sundaresan's divergence arise as optimal solutions. This motivates us to establish a connection among these four problems. In this paper, we study a more general problem and show that Rényi and Shannon entropies arise as its solution. We show that the problems on source coding, guessing and task partitioning are particular instances of this general optimization problem, and derive the lower bounds using this framework. We also refine some known results and present new results for mismatched version of these problems using a unified approach. We strongly feel that this generalization would, in addition to help in understanding the similarities and distinctiveness of these problems, also help to solve any new problem that falls in this framework.

cs.IT

Generalized Estimating Equation for the Student-t Distributions

In \cite{KumarS15J2}, it was shown that a generalized maximum likelihood estimation problem on a (canonical) $α$-power-law model ($\mathbb{M}^{(α)}$-family) can be solved by solving a system of linear equations. This was due to an orthogonality relationship between the $\mathbb{M}^{(α)}$-family and a linear family with respect to the relative $α$-entropy (or the $\mathscr{I}_α$-divergence). Relative $α$-entropy is a generalization of the usual relative entropy (or the Kullback-Leibler divergence). $\mathbb{M}^{(α)}$-family is a generalization of the usual exponential family. In this paper, we first generalize the $\mathbb{M}^{(α)}$-family including the multivariate, continuous case and show that the Student-t distributions fall in this family. We then extend the above stated result of \cite{KumarS15J2} to the general $\mathbb{M}^{(α)}$-family. Finally we apply this result to the Student-t distribution and find generalized estimators for its parameters.

math.ST

Information Geometric Approach to Bayesian Lower Error Bounds

Information geometry describes a framework where probability densities can be viewed as differential geometry structures. This approach has shown that the geometry in the space of probability distributions that are parameterized by their covariance matrix is linked to the fundamentals concepts of estimation theory. In particular, prior work proposes a Riemannian metric - the distance between the parameterized probability distributions - that is equivalent to the Fisher Information Matrix, and helpful in obtaining the deterministic Cramér-Rao lower bound (CRLB). Recent work in this framework has led to establishing links with several practical applications. However, classical CRLB is useful only for unbiased estimators and inaccurately predicts the mean square error in low signal-to-noise (SNR) scenarios. In this paper, we propose a general Riemannian metric that, at once, is used to obtain both Bayesian CRLB and deterministic CRLB along with their vector parameter extensions. We also extend our results to the Barankin bound, thereby enhancing their applicability to low SNR situations.

cs.IT

Projection Theorems of Divergences and Likelihood Maximization Methods

Projection theorems of divergences enable us to find reverse projection of a divergence on a specific statistical model as a forward projection of the divergence on a different but rather "simpler" statistical model, which, in turn, results in solving a system of linear equations. Reverse projection of divergences are closely related to various estimation methods such as the maximum likelihood estimation or its variants in robust statistics. We consider projection theorems of three parametric families of divergences that are widely used in robust statistics, namely the Rényi divergences (or the Cressie-Reed power divergences), density power divergences, and the relative $α$-entropy (or the logarithmic density power divergences). We explore these projection theorems from the usual likelihood maximization approach and from the principle of sufficiency. In particular, we show the equivalence of solving the estimation problems by the projection theorems of the respective divergences and by directly solving the corresponding estimating equations. We also derive the projection theorem for the density power divergences.

cs.IT

Projection Theorems for the Rényi Divergence on $α$-Convex Sets

This paper studies forward and reverse projections for the Rényi divergence of order $α\in (0, \infty)$ on $α$-convex sets. The forward projection on such a set is motivated by some works of Tsallis {\em et al.} in statistical physics, and the reverse projection is motivated by robust statistics. In a recent work, van Erven and Harremoës proved a Pythagorean inequality for Rényi divergences on $α$-convex sets under the assumption that the forward projection exists. Continuing this study, a sufficient condition for the existence of forward projection is proved for probability measures on a general alphabet. For $α\in (1, \infty)$, the proof relies on a new Apollonius theorem for the Hellinger divergence, and for $α\in (0,1)$, the proof relies on the Banach-Alaoglu theorem from functional analysis. Further projection results are then obtained in the finite alphabet setting. These include a projection theorem on a specific $α$-convex set, which is termed an {\em $α$-linear family}, generalizing a result by Csiszár for $α\neq 1$. The solution to this problem yields a parametric family of probability measures which turns out to be an extension of the exponential family, and it is termed an {\em $α$-exponential family}. An orthogonality relationship between the $α$-exponential and $α$-linear families is established, and it is used to turn the reverse projection on an $α$-exponential family into a forward projection on a $α$-linear family. This paper also proves a convergence result of an iterative procedure used to calculate the forward projection on an intersection of a finite number of $α$-linear families.

cs.IT

Minimization Problems Based on Relative $α$-Entropy I: Forward Projection

Minimization problems with respect to a one-parameter family of generalized relative entropies are studied. These relative entropies, which we term relative $α$-entropies (denoted $\mathscr{I}_α$), arise as redundancies under mismatched compression when cumulants of compressed lengths are considered instead of expected compressed lengths. These parametric relative entropies are a generalization of the usual relative entropy (Kullback-Leibler divergence). Just like relative entropy, these relative $α$-entropies behave like squared Euclidean distance and satisfy the Pythagorean property. Minimizers of these relative $α$-entropies on closed and convex sets are shown to exist. Such minimizations generalize the maximum Rényi or Tsallis entropy principle. The minimizing probability distribution (termed forward $\mathscr{I}_α$-projection) for a linear family is shown to obey a power-law. Other results in connection with statistical inference, namely subspace transitivity and iterated projections, are also established. In a companion paper, a related minimization problem of interest in robust statistics that leads to a reverse $\mathscr{I}_α$-projection is studied.

cs.IT

Minimization Problems Based on Relative $α$-Entropy II: Reverse Projection

In part I of this two-part work, certain minimization problems based on a parametric family of relative entropies (denoted $\mathscr{I}_α$) were studied. Such minimizers were called forward $\mathscr{I}_α$-projections. Here, a complementary class of minimization problems leading to the so-called reverse $\mathscr{I}_α$-projections are studied. Reverse $\mathscr{I}_α$-projections, particularly on log-convex or power-law families, are of interest in robust estimation problems ($α>1$) and in constrained compression settings ($α<1$). Orthogonality of the power-law family with an associated linear family is first established and is then exploited to turn a reverse $\mathscr{I}_α$-projection into a forward $\mathscr{I}_α$-projection. The transformed problem is a simpler quasiconvex minimization subject to linear constraints.

cs.IT

Relative $α$-Entropy Minimizers Subject to Linear Statistical Constraints

We study minimization of a parametric family of relative entropies, termed relative $α$-entropies (denoted $\mathscr{I}_α(P,Q)$). These arise as redundancies under mismatched compression when cumulants of compressed lengths are considered instead of expected compressed lengths. These parametric relative entropies are a generalization of the usual relative entropy (Kullback-Leibler divergence). Just like relative entropy, these relative $α$-entropies behave like squared Euclidean distance and satisfy the Pythagorean property. Minimization of $\mathscr{I}_α(P,Q)$ over the first argument on a set of probability distributions that constitutes a linear family is studied. Such a minimization generalizes the maximum Rényi or Tsallis entropy principle. The minimizing probability distribution (termed $\mathscr{I}_α$-projection) for a linear family is shown to have a power-law.

cs.IT