SearcharxivSearch

arXiv subjects

Neri Merhav

Publications and source records attributed to Neri Merhav.

At least 19 recordsLinked to original sources

Universal Random Coding for Successive Refinement of Individual Sequences Based on Lempel-Ziv Complexity

This work extends an earlier proposed universal random-coding ensemble for sample-wise lossy compression of individual sequences to the two-stage (successive-refinement) setting --- an extension that is not quite straightforward, as explained in the sequel. The construction has three layers. First, a complete, self-contained direct/converse match in the language of type classes and typical sequences, a two-stage analogue of Rimoldi's classical characterization (applied in the superalphabet of blocks). Second, a realization of the same rates via a random-coding scheme built from the Lempel--Ziv (LZ78) complexity rather than the type-class uniform measure the first layer relies on, since an LZ-based ensemble is universal across source statistics and distortion measures, carries no block-length parameter in the codebook itself, and is directly implementable. Third, we give a fully type-free achievability construction (no joint types of $\ell$-vectors appear anywhere in its search criterion) and show it matches the same converse exactly, reaching every point of the achievable region: the same converse from the first layer already binds any code, type-free schemes included, so no separate argument is needed to certify the match. We also describe (without proof) how to extend the type-free construction to any fixed number $r>2$ of stages. Finally, we show this paper's achievable region is a superset of the one obtained by a companion paper, which is based on a finite-state modeling approach to a similar two-stage problem, extending a known single-stage domination result to two stages.

cs.IT

Statistics of the Compression Ratio of a Variable-to-Variable Code: Exact Moments and Asymptotic Behavior

A variable-to-variable (V2V) length code parses a source sequence into phrases of variable length and maps each phrase to a binary codeword of, generally, a different random length. After encoding $n$ phrases, the realized compression ratio $R_n=\Lambda_n/\Sigma_n$ -- total codeword length over total source-symbol count -- is the finite-sample counterpart of the code's asymptotic rate $\rho$, to which it converges only as $n\to\infty$. This paper first derives exact formulas for all integer moments of $R_n$ for a given discrete memoryless source (DMS). Specifically, we obtain a closed-form formula for every moment $\E\{R_n^k\}$ as a one-dimensional integral involving only single-phrase moment generating functions of the pair $(L,\ell)$ -- the phrase length, in source symbols, and codeword length, in bits. From these moments we derive an Edgeworth approximation to the cumulative distribution function (CDF) of $R_n$ that is substantially more accurate than the central limit theorem (CLT) approximation. Using the Laplace method of integration, we also derive explicit closed-form formulas for the bias constant $C=\lim_{n\to\infty}n(\E\{R_n\}-\rho)$ and for the variance constant $\lim_{n\to\infty}n\cdot\Var\{R_n\}$. The analysis extends to Markov sources via state-indexed matrices with a redundancy formula obtained in closed form. On the coding-theoretic side, we cast V2V length codes as finite-state encoders and apply a generalized Kraft inequality for a compression-rate lower bound, and give a structural decomposition of the bias coefficient that separates cleanly across variable-to-fixed (V2F) length codes, fixed-to-variable (F2V) length codes, and V2V length codes. Applied to the Khodak code of Bugeaud, Drmota, and Szpankowski, this decomposition shows that its improved performance is reflected in its smaller bias constant.

cs.IT

Grouped Reverse Importance Sampling for the Partition Function

We introduce and analyze several grouped variants of the method of reverse importance sampling (RIS) for estimating a partition function from samples of the Boltzmann distribution $p(x)=e^{ \betaU(x)}/Z(\beta)$. Ordinary RIS weighs each sample separately. By contrast, our proposed grouped RIS (GRIS) methods are based on assigning the samples into groups (or batches) of size $k\ge 2$ and applying a joint weight function to each group. The focal point of the research is the quest for a tractable weight function that would yield the smallest possible mean squared error (MSE). A simple identity relates the normalized MSE to the chi-squared divergence between the joint-weight distribution and the distribution of the $k$-fold sum of independent energies. Our first theoretical finding is that any weight that improves on ordinary RIS ($k=1$) must couple the group components. In other words, it must not be a product-form function across those components, as product-form weight functions always worsen the MSE. Our second, and more important, finding is that, without loss of optimality, it is sufficient to seek weight functions that depend only on the total energy, $\sum_iU(x_i)$, of the group (group-energy weight functions); for the sliding-window variants, the analogous result is open. This finding simplifies both the theoretical analysis and the application of the method substantially. For $k=2$ and $k=3$, the MSE associated with non-overlapping (NOL) groups is reduced by $20$--$65\%$ across three examples. We then propose two additional variants of GRIS, both based on sliding-window grouping (as opposed to NOL grouping). The first applies a fixed weight sliding window (FSW) across all (cyclic) shifts of the sliding window, and the second allows a variable-weight sliding window (VSW). The FSW scheme improves on the NOL one, and the VSW improves even further, as will be demonstrated numerically.

cs.IT

Soft Covering via Hypothesis Testing: Typical-Code Exponents and Mismatched Detection

We study the typical-code (quenched) behavior of the false-alarm (FA) and missed-detection (MD) error exponents of the Neyman-Pearson test associated with soft covering, complementing the average-code (annealed) analysis that has been carried out in a companion paper [1]. We prove that, as the block-length tends to infinity, for almost every randomly selected fixed-composition codebook, the negative normalized logarithms of both error probabilities converge to their respective average-code exponents. In other words, the error exponents are self-averaging. We then extend the scope and study a mismatched likelihood ratio test that assumes the wrong channel model. Here, we derive the mismatched error exponents, show that self-averaging persists under mismatch, and characterize the degradation. In particular, we characterize the coding rate beyond which the two kinds of error exponents cannot be positive at the same time, which in the matched case, is given by the channel input-output mutual information rate.

cs.IT

Soft Covering Through the Lens of Hypothesis Testing

We study the soft covering phenomenon through the lens of Neyman--Pearson hypothesis testing: given a channel output sequence $y^n$, can one decide whether it was produced when the channel was driven by a random codeword, or generated independently from the output marginal? We derive exact exponential decay rates for the jointly averaged false-alarm (FA) probability $\alpha_n(\tau,R)$ and missed-detection (MD) probability $\beta_n(\tau,R)$, as functions of the decision threshold $\tau$ and the codebook rate $R$. The derived single-letter formulas of the exponents $\EFA(\tau,R)=-\lim_{n\to\infty}\frac{1}{n}\ln\alpha_n(\tau,R)$ and $\EMD(\tau,R)=-\lim_{n\to\infty}\frac{1}{n}\ln\beta_n(\tau,R)$ are tight in the random coding sense. The analysis reveals a rich phase structure. For $R < I(X;Y)$, there is a genuine exponential tradeoff between the two error types over the interval $\tau \in (0, I(X;Y)-R)$. At $R = I(X;Y)$, this tradeoff interval collapses to the single point $\tau = 0$, where both error exponents simultaneously vanish, a fact which manifests the soft covering phenomenon in the Neyman--Pearson sense. For $R > I(X;Y)$, the same instantaneous collapse persists at $\tau = 0$; moreover, for every $\tau$ at least one exponent is zero: the FA exponent is zero for $\tau \le 0$ (FA probability does not decay exponentially), and the MD exponent is zero for $\tau \ge 0$ (and finite, channel-specific for $\tau<0$; see Remark~\ref{rem:jump}). There is no interval of $\tau$ where both exponents are simultaneously positive. A sharp phase transition in the MD exponent occurs at $\tau^* = [I(X;Y)-R]_+$ for all rates.

cs.IT

A Statistical-Physics Refinement of Soft Covering

We study the channel output distribution induced by a rate-$R$ random code via statistical physics. The partition function is $Z_n(\beta|\mathcal{C}) = \sum_{y^n}[P_{Y^n|\mathcal{C}}(y^n)]^\beta$, where $\mathcal{C}$ is the code and $\beta > 0$ is inverse temperature. Our focus is on the free energy which is the normalized logarithm of this quantity, which encodes the full R\'{e}nyi spectrum of the output distribution. The single-letter formula derived for the annealed free energy decomposes into two branches which reflect a ``competition'' between two populations of codewords. One is the \emph{bulk branch}, $\psi_{\mbox{\tiny b}}(\beta,R)$, which is driven by typical codewords and the other one is the \emph{sparse branch} $\psi_{\mbox{\tiny s}}(\beta,R)$, which is driven by a-typical codewords, where the qualifiers `typical' and `atypical' are in a sense to become apparent later. We analyze the phase structure of each branch separately and characterize their competition. Both branches are derived for all $\beta > 0$. The phase boundary $R^\star(\beta)$, where the two branches are equal, is analyzed for $\beta \geq 1$, where it has an explicit closed-form expression. The phase diagram in the first quadrant of the $(\beta, R)$ plane has four regions separated by three boundaries: $R = I^{\mbox{\tiny b}}(\beta)$ (bulk branch transition), $R = R^\star(\beta)$ (bulk--sparse competition boundary), and $R = I^{\mbox{\tiny s}}(\beta)$ (sparse branch transition), all meeting at the point $(\beta, R) = (1, I(X;Y))$, where $I(X;Y)$ is the mutual information induced by the input type and the channel. Applications to guesswork, channel resolvability, and hypothesis testing are discussed, and all results are illustrated with a numerical example of a Z-channel.

cs.IT

Statistical Physics of Coding for the Integers

We study a paradigm of coding for compression of the natural numbers via the zeta distribution and develop a statistical-mechanical interpretation, both in terms of Hagedorn systems and a Bose gas with energy levels given by logarithms of prime numbers. We also propose a simple coding scheme for the zeta distribution that nearly achieves the ideal code length. For block coding of vectors of natural numbers, we derive the micro-canonical entropy function and demonstrate its asymptotic linearity implying that its behavior is analogous to that of a Hagedorn system. We also derive the large deviations rate function, and provide a formula for the best coding parameter in the large deviations sense. We show that due the Hagedorn-type phase transition there is only partial equivalence of ensembles, due to the degeneration of the domain of the partition function.

cond-mat.stat-mech

Generalized Forms of the Kraft Inequality for Finite-State Encoders

We derive a few extended versions of the Kraft inequality for information lossless finite-state encoders. The main basic contribution is in defining a notion of a Kraft matrix and in establishing the fact that a necessary condition for information losslessness of a finite-state encoder is that none of the eigenvalues of this matrix have modulus larger than unity, or equivalently, the generalized Kraft inequality asserts that the spectral radius of the Kraft matrix cannot exceed one. For the important special case where the FS encoder is irreducible, we derive several equivalent forms of this inequality, which are based on well known formulas for spectral radius. It also turns out that in the irreducible case, Kraft sums are bounded by a constant, independent of the block length, and thus cannot grow even in any subexponential rate. Finally, two extensions are outlined - one concerns the case of side information available to both encoder and decoder, and the other is for lossy compression.

cs.IT

Refinements and Generalizations of the Shannon Lower Bound via Extensions of the Kraft Inequality

We derive a few extended versions of the Kraft inequality for lossy compression, which pave the way to the derivation of several refinements and extensions of the well known Shannon lower bound in a variety of instances of rate-distortion coding. These refinements and extensions include sharper bounds for one-to-one codes and $D$-semifaithful codes, a Shannon lower bound for distortion measures based on sliding-window functions, and an individual-sequence counterpart of the Shannon lower bound.

cs.IT

Volume-Based Lower Bounds to the Capacity of the Gaussian Channel Under Pointwise Additive Input Constraints

We present a family of relatively simple and unified lower bounds on the capacity of the Gaussian channel under a set of pointwise additive input constraints. Specifically, the admissible channel input vectors $\bx = (x_1, \ldots, x_n)$ must satisfy $k$ additive cost constraints of the form $\sum_{i=1}^n \phi_j(x_i) \le n \Gamma_j$, $j = 1,2,\ldots,k$, which are enforced pointwise for every $\bx$, rather than merely in expectation. More generally, we also consider cost functions that depend on a sliding window of fixed length $m$, namely, $\sum_{i=m}^n \phi_j(x_i, x_{i-1}, \ldots, x_{i-m+1}) \le n \Gamma_j$, $j = 1,2,\ldots,k$, a formulation that naturally accommodates correlation constraints as well as a broad range of other constraints of practical relevance. We propose two classes of lower bounds, derived by two methodologies that both rely on the exact evaluation of the volume exponent associated with the set of input vectors satisfying the given constraints. This evaluation exploits extensions of the method of types to continuous alphabets, the saddle-point method of integration, and basic tools from large deviations theory. The first class of bounds is obtained via the entropy power inequality (EPI), and therefore applies exclusively to continuous-valued inputs. The second class, by contrast, is more general, and it applies to discrete input alphabets as well. It is based on a direct manipulation of mutual information, and it yields stronger and tighter bounds, though at the cost of greater technical complexity. Numerical examples illustrating both types of bounds are provided, and several extensions and refinements are also discussed.

cs.IT

Lempel-Ziv Complexity, Empirical Entropies, and Chain Rules

We derive upper and lower bounds on the overall compression ratio of the 1978 Lempel-Ziv (LZ78) algorithm, applied independently to $k$-blocks of a finite individual sequence. Both bounds are given in terms of normalized empirical entropies of the given sequence. For the bounds to be tight and meaningful, the order of the empirical entropy should be small relative to $k$ in the upper bound, but large relative to $k$ in the lower bound. Several non-trivial conclusions arise from these bounds. One of them is a certain form of a chain rule of the Lempel-Ziv (LZ) complexity, which decomposes the joint LZ complexity of two sequences, say, $\bx$ and $\by$, into the sum of the LZ complexity of $\bx$ and the conditional LZ complexity of $\by$ given $\bx$ (up to small terms). The price of this decomposition, however, is in changing the length of the block. Additional conclusions are discussed as well.

cs.IT

Universal Encryption of Individual Sequences Under Maximal Leakage

We consider the Shannon cipher system in the framework of individual sequences and finite-state encrypters under the metric of maximal leakage of information. A lower bound and an asymptotically matching upper bound on the leakage are derived, which lead to the conclusion that asymptotically minimum leakage can be attained by Lempel-Ziv compression followed by one-time pad encryption of the compressed bit-stream.

cs.IT

Successive Refinement for Lossy Compression of Individual Sequences

We consider the problem of successive-refinement coding for lossy compression of individual sequences, namely, compression in two stages, where in the first stage, a coarse description at a relatively low rate is sent from the encoder to the decoder, and in the second stage, additional coding rate is allocated in order to refine the description and thereby improve the reproduction. Our main result is in establishing outer bounds (converse theorems) for the rate region where we limit the encoders to be finite-state machines in the spirit of Ziv and Lempel's 1978 model.The matching achievability scheme is conceptually straightforward. We also consider the more general multiple description coding problem on a similar footing and propose achievability schemes that are analogous to the well-known El Gamal-Cover and the Zhang-Berger achievability schemes of memoryless sources and additive distortion measures.

cs.IT

Two New Families of Local Asymptotically Minimax Lower Bounds in Parameter Estimation

We propose two families of asymptotically local minimax lower bounds on parameter estimation performance. The first family of bounds applies to any convex, symmetric loss function that depends solely on the difference between the estimate and the true underlying parameter value (i.e., the estimation error), whereas the second is more specifically oriented to the moments of the estimation error. The proposed bounds are relatively easy to calculate numerically (in the sense that their optimization is over relatively few auxiliary parameters), yet they turn out to be tighter (sometimes significantly so) than previously reported bounds that are associated with similar calculation efforts, across a variety of application examples. In addition to their relative simplicity, they also have the following advantages: (i) Essentially no regularity conditions are required regarding the parametric family of distributions; (ii) The bounds are local (in a sense to be specified); (iii) The bounds provide the correct order of decay as functions of the number of observations, at least in all examples examined; (iv) At least the first family of bounds extends straightforwardly to vector parameters.

math.ST

On Jacob Ziv's Individual-Sequence Approach to Information Theory

This article stands as a tribute to the enduring legacy of Jacob Ziv and his landmark contributions to information theory. Specifically, it delves into the groundbreaking individual-sequence approach -- a cornerstone of Ziv's academic pursuits. Together with Abraham Lempel, Ziv pioneered the renowned Lempel-Ziv (LZ) algorithm, a beacon of innovation in various versions. Beyond its original domain of universal data compression, this article underscores the broad utility of the individual-sequence approach and the LZ algorithm across a wide spectrum of problem areas. As we traverse through the forthcoming pages, it will also become evident how Ziv's visionary approach has left an indelible mark on my own research journey, as well as on those of numerous colleagues and former students. We shall explore, not only the technical power of the LZ algorithm, but also its profound impact on shaping the landscape of information theory and its applications.

cs.IT

A Toolbox for Refined Information-Theoretic Analyses with Applications

This monograph offers a toolbox of mathematical techniques, which have been effective and widely applicable in information-theoretic analysis. The first tool is a generalization of the method of types to Gaussian settings, and then to general exponential families. The second tool is Laplace and saddle-point integration, which allow to refine the results of the method of types, and are capable of obtaining more precise results. The third is the type class enumeration method, a principled method to evaluate the exact random-coding exponent of coded systems, which results in the best known exponent in various problem settings. The fourth subset of tools aimed at evaluating the expectation of non-linear functions of random variables, either via integral representations, or by a refinement of Jensen's inequality via change-of-measure, by complementing Jensen's inequality with a reversed inequality, or by a class of generalized Jensen's inequalities that are applicable for functions beyond convex/concave. Various application examples of all these tools are provided along this monograph.

cs.IT

Refinements and Extensions of Ziv's Model of Perfect Secrecy for Individual Sequences

We refine and extend Ziv's model and results regarding perfectly secure encryption of individual sequences. According to this model, the encrypter and the legitimate decrypter share in common a secret key, not shared with the unauthorized eavesdropper, who is aware of the encryption scheme and has some prior knowledge concerning the individual plaintext source sequence. This prior knowledge, combined with the cryptogram, is harnessed by eavesdropper which implements a finite-state machine as a mechanism for accepting or rejecting attempted guesses of the source plaintext. The encryption is considered perfectly secure if the cryptogram does not provide any new information to the eavesdropper that may enhance its knowledge concerning the plaintext beyond his prior knowledge. Ziv has shown that the key rate needed for perfect secrecy is essentially lower bounded by the finite-state compressibility of the plaintext sequence, a bound which is clearly asymptotically attained by Lempel-Ziv compression followed by one-time pad encryption. In this work, we consider some more general classes of finite-state eavesdroppers and derive the respective lower bounds on the key rates needed for perfect secrecy. These bounds are tighter and more refined than Ziv's bound and they are attained by encryption schemes that are based on different universal lossless compression schemes. We also extend our findings to the case where side information is available to the eavesdropper and the legitimate decrypter, but may or may not be available to the encrypter as well.

cs.IT

Optimal Signals and Detectors Based on Correlation and Energy

In continuation of an earlier study, we explore a Neymann-Pearson hypothesis testing scenario where, under the null hypothesis ($\cal{H}_0$), the received signal is a white noise process $N_t$, which is not Gaussian in general, and under the alternative hypothesis ($\cal{H}_1$), the received signal comprises a deterministic transmitted signal $s_t$ corrupted by additive white noise, the sum of $N_t$ and another noise process originating from the transmitter, denoted as $Z_t$, which is not necessarily Gaussian either. Our approach focuses on detectors that are based on the correlation and energy of the received signal, which are motivated by implementation simplicity. We optimize the detector parameters to achieve the best trade-off between missed-detection and false-alarm error exponents. First, we optimize the detectors for a given signal, resulting in a non-linear relation between the signal and correlator weights to be optimized. Subsequently, we optimize the transmitted signal and the detector parameters jointly, revealing that the optimal signal is a balanced ternary signal and the correlator has at most three different coefficients, thus facilitating a computationally feasible solution.

cs.IT