SearcharxivSearch

arXiv subjects

N. Santhanam

Publications and source records attributed to N. Santhanam.

2 recordsLinked to original sources

Data driven weak universal consistency

Many current applications in data science need rich model classes to adequately represent the statistics that may be driving the observations. But rich model classes may be too complex to admit estimators that converge to the truth with convergence rates that can be uniformly bounded over the entire collection of probability distributions comprising the model class, i.e. it may be impossible to guarantee uniform consistency of such estimators as the sample size increases. In such cases, it is conventional to settle for estimators with guarantees on convergence rate where the performance can be bounded in a model-dependent way, i.e. pointwise consistent estimators. But this viewpoint has the serious drawback that estimator performance is a function of the unknown model within the model class that is being estimated, and is therefore unknown. Even if an estimator is consistent, how well it is doing at any given time may not be clear, no matter what the sample size of the observations. Departing from the classical uniform/pointwise consistency dichotomy that leads to this impasse, a new analysis framework is explored by studying rich model classes that may only admit pointwise consistency guarantees, yet all the information about the unknown model driving the observations that is needed to gauge estimator accuracy can be inferred from the sample at hand. We expect that this data-derived estimation framework will be broadly applicable to a wide range of estimation problems by providing a methodology to deal with much richer model classes. In this paper we analyze the lossless compression problem in detail in this novel data-derived framework.

cs.IT

Universal compression of Gaussian sources with unknown parameters

For a collection of distributions over a countable support set, the worst case universal compression formulation by Shtarkov attempts to assign a universal distribution over the support set. The formulation aims to ensure that the universal distribution does not underestimate the probability of any element in the support set relative to distributions in the collection. When the alphabet is uncountable and we have a collection $\cal P$ of Lebesgue continuous measures instead, we ask if there is a corresponding universal probability density function (pdf) that does not underestimate the value of the density function at any point in the support relative to pdfs in $\cal P$. Analogous to the worst case redundancy of a collection of distributions over a countable alphabet, we define the \textit{attenuation} of a class to be $A$ when the worst case optimal universal pdf at any point $x$ in the support is always at least the value any pdf in the collection $\cal P$ assigns to $x$ divided by $A$. We analyze the attenuation of the worst optimal universal pdf over length-$n$ samples generated \textit{i.i.d.} from a Gaussian distribution whose mean can be anywhere between $-α/2$ to $α/2$ and variance between $σ_m^2$ and $σ_M^2$. We show that this attenuation is finite, grows with the number of samples as ${\cal O}(n)$, and also specify the attentuation exactly without approximations. When only one parameter is allowed to vary, we show that the attenuation grows as ${\cal O}(\sqrt{n})$, again keeping in line with results from prior literature that fix the order of magnitude as a factor of $\sqrt{n}$ per parameter. In addition, we also specify the attenuation exactly without approximation when only the mean or only the variance is allowed to vary.

cs.IT