Searcharxiv⌕ Search

arXiv subjects

Alex Dytso

Publications and source records attributed to Alex Dytso.

At least 37 records · Page 2Linked to original sources

On $2 \times 2$ MIMO Gaussian Channels with a Small Discrete-Time Peak-Power Constraint

A multi-input multi-output (MIMO) Gaussian channel with two transmit antennas and two receive antennas is studied that is subject to an input peak-power constraint. The capacity and the capacity-achieving input distribution are unknown in general. The problem is shown to be equivalent to a channel with an identity matrix but where the input lies inside and on an ellipse with principal axis length $r_p$ and minor axis length $r_m$. If $r_p \le \sqrt{2}$, then the capacity-achieving input has support on the ellipse. A sufficient condition is derived under which a two-point distribution is optimal. Finally, if $r_m < r_p \le \sqrt{2}$, then the capacity-achieving distribution is discrete.

cs.IT↗

Data-Driven Estimation of the False Positive Rate of the Bayes Binary Classifier via Soft Labels

Classification is a fundamental task in many applications on which data-driven methods have shown outstanding performances. However, it is challenging to determine whether such methods have achieved the optimal performance. This is mainly because the best achievable performance is typically unknown and hence, effectively estimating it is of prime importance. In this paper, we consider binary classification problems and we propose an estimator for the false positive rate (FPR) of the Bayes classifier, that is, the optimal classifier with respect to accuracy, from a given dataset. Our method utilizes soft labels, or real-valued labels, which are gaining significant traction thanks to their properties. We thoroughly examine various theoretical properties of our estimator, including its consistency, unbiasedness, rate of convergence, and variance. To enhance the versatility of our estimator beyond soft labels, we also consider noisy labels, which encompass binary labels. For noisy labels, we develop effective FPR estimators by leveraging a denoising technique and the Nadaraya-Watson estimator. Due to the symmetry of the problem, our results can be readily applied to estimate the false negative rate of the Bayes classifier.

cs.LG↗

Binomial Channel: On the Capacity-Achieving Distribution and Bounds on the Capacity

This work considers a binomial noise channel. The paper can be roughly divided into two parts. The first part is concerned with the properties of the capacity-achieving distribution. In particular, for the binomial channel, it is not known if the capacity-achieving distribution is unique since the output space is finite (i.e., supported on integers $0, \ldots, n)$ and the input space is infinite (i.e., supported on the interval $[0,1]$), and there are multiple distributions that induce the same output distribution. This paper shows that the capacity-achieving distribution is unique by appealing to the total positivity property of the binomial kernel. In addition, we provide upper and lower bounds on the cardinality of the support of the capacity-achieving distribution. Specifically, an upper bound of order $ \frac{n}{2}$ is shown, which improves on the previous upper bound of order $n$ due to Witsenhausen. Moreover, a lower bound of order $\sqrt{n}$ is shown. Finally, additional information about the locations and probability values of the support points is established. The second part of the paper focuses on deriving upper and lower bounds on capacity. In particular, firm bounds are established for all $n$ that show that the capacity scales as $\frac{1}{2} \log(n)$.

cs.IT↗

Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning

We consider information-theoretic bounds on expected generalization error for statistical learning problems in a networked setting. In this setting, there are $K$ nodes, each with its own independent dataset, and the models from each node have to be aggregated into a final centralized model. We consider both simple averaging of the models as well as more complicated multi-round algorithms. We give upper bounds on the expected generalization error for a variety of problems, such as those with Bregman divergence or Lipschitz continuous losses, that demonstrate an improved dependence of $1/K$ on the number of nodes. These "per node" bounds are in terms of the mutual information between the training dataset and the trained weights at each node, and are therefore useful in describing the generalization properties inherent to having communication or privacy constraints at each node.

cs.IT↗

Improved Bounds on the Number of Support Points of the Capacity-Achieving Input for Amplitude Constrained Poisson Channels

This work considers a discrete-time Poisson noise channel with an input amplitude constraint $\mathsf{A}$ and a dark current parameter $λ$. It is known that the capacity-achieving distribution for this channel is discrete with finitely many points. Recently, for $λ=0$, a lower bound of order $\sqrt{\mathsf{A}}$ and an upper bound of order $\mathsf{A} \log^2(\mathsf{A})$ have been demonstrated on the cardinality of the support of the optimal input distribution. In this work, we improve these results in several ways. First, we provide upper and lower bounds that hold for non-zero dark current. Second, we produce a sharper upper bound with a far simpler technique. In particular, for $λ=0$, we sharpen the upper bound from the order of $\mathsf{A} \log^2(\mathsf{A})$ to the order of $\mathsf{A}$. Finally, some other additional information about the location of the support is provided.

cs.IT↗

Uniform Distribution on $(n-1)$-Sphere: Rate-Distortion under Squared Error Distortion

This paper investigates the rate-distortion function, under a squared error distortion $D$, for an $n$-dimensional random vector uniformly distributed on an $(n-1)$-sphere of radius $R$. First, an expression for the rate-distortion function is derived for any values of $n$, $D$, and $R$. Second, two types of asymptotics with respect to the rate-distortion function of a Gaussian source are characterized. More specifically, these asymptotics concern the low-distortion regime (that is, $D \to 0$) and the high-dimensional regime (that is, $n \to \infty$).

cs.IT↗

Functional Properties of the Ziv-Zakai bound with Arbitrary Inputs

This paper explores the Ziv-Zakai bound (ZZB), which is a well-known Bayesian lower bound on the Minimum Mean Squared Error (MMSE). First, it is shown that the ZZB holds without any assumption on the distribution of the estimand, that is, the estimand does not necessarily need to have a probability density function. The ZZB is then further analyzed in the high-noise and low-noise regimes and shown to always tensorize. Finally, the tightness of the ZZB is investigated under several aspects, such as the number of hypotheses and the usefulness of the valley-filling function. In particular, a sufficient and necessary condition for the tightness of the bound with continuous inputs is provided, and it is shown that the bound is never tight for discrete input distributions with a support set that does not have an accumulation point at zero.

cs.IT↗

Amplitude Constrained Vector Gaussian Wiretap Channel: Properties of the Secrecy-Capacity-Achieving Input Distribution

This paper studies secrecy-capacity of an $n$-dimensional Gaussian wiretap channel under a peak-power constraint. This work determines the largest peak-power constraint $\bar{\mathsf{R}}_n$ such that an input distribution uniformly distributed on a single sphere is optimal; this regime is termed the low amplitude regime. The asymptotic of $\bar{\mathsf{R}}_n$ as $n$ goes to infinity is completely characterized as a function of noise variance at both receivers. Moreover, the secrecy-capacity is also characterized in a form amenable for computation. Several numerical examples are provided, such as the example of the secrecy-capacity-achieving distribution beyond the low amplitude regime. Furthermore, for the scalar case $(n=1)$ we show that the secrecy-capacity-achieving input distribution is discrete with finitely many points at most of the order of $\frac{\mathsf{R}^2}{σ_1^2}$, where $σ_1^2$ is the variance of the Gaussian noise over the legitimate channel.

cs.IT↗

An MMSE Lower Bound via Poincaré Inequality

This paper studies the minimum mean squared error (MMSE) of estimating $\mathbf{X} \in \mathbb{R}^d$ from the noisy observation $\mathbf{Y} \in \mathbb{R}^k$, under the assumption that the noise (i.e., $\mathbf{Y}|\mathbf{X}$) is a member of the exponential family. The paper provides a new lower bound on the MMSE. Towards this end, an alternative representation of the MMSE is first presented, which is argued to be useful in deriving closed-form expressions for the MMSE. This new representation is then used together with the Poincaré inequality to provide a new lower bound on the MMSE. Unlike, for example, the Cramér-Rao bound, the new bound holds for all possible distributions on the input $\mathbf{X}$. Moreover, the lower bound is shown to be tight in the high-noise regime for the Gaussian noise setting under the assumption that $\mathbf{X}$ is sub-Gaussian. Finally, several numerical examples are shown which demonstrate that the bound performs well in all noise regimes.

cs.IT↗

Entropic CLT for Order Statistics

It is well known that central order statistics exhibit a central limit behavior and converge to a Gaussian distribution as the sample size grows. This paper strengthens this known result by establishing an entropic version of the CLT that ensures a stronger mode of convergence using the relative entropy. In particular, an order $O(1/\sqrt{n})$ rate of convergence is established under mild conditions on the parent distribution of the sample generating the order statistics. To prove this result, ancillary results on order statistics are derived, which might be of independent interest.

cs.IT↗

A Dimensionality Reduction Method for Finding Least Favorable Priors with a Focus on Bregman Divergence

A common way of characterizing minimax estimators in point estimation is by moving the problem into the Bayesian estimation domain and finding a least favorable prior distribution. The Bayesian estimator induced by a least favorable prior, under mild conditions, is then known to be minimax. However, finding least favorable distributions can be challenging due to inherent optimization over the space of probability distributions, which is infinite-dimensional. This paper develops a dimensionality reduction method that allows us to move the optimization to a finite-dimensional setting with an explicit bound on the dimension. The benefit of this dimensionality reduction is that it permits the use of popular algorithms such as projected gradient ascent to find least favorable priors. Throughout the paper, in order to make progress on the problem, we restrict ourselves to Bayesian risks induced by a relatively large class of loss functions, namely Bregman divergences.

stat.ML↗

On the Capacity Achieving Input of Amplitude Constrained Vector Gaussian Wiretap Channel

This paper studies secrecy-capacity of an $n$-dimensional Gaussian wiretap channel under the peak-power constraint. This work determines the largest peak-power constraint $\bar{\mathsf{R}}_n$ such that an input distribution uniformly distributed on a single sphere is optimal; this regime is termed the small-amplitude regime. The asymptotic of $\bar{\mathsf{R}}_n$ as $n$ goes to infinity is completely characterized as a function of noise variance at both receivers. Moreover, the secrecy-capacity is also characterized in a form amenable for computation. Furthermore, several numerical examples are provided, such as the example of the secrecy-capacity achieving distribution outside of the small amplitude regime.

cs.IT↗

Poisson Noise Channel with Dark Current: Numerical Computation of the Optimal Input Distribution

This paper considers a discrete time-Poisson noise channel which is used to model pulse-amplitude modulated optical communication with a direct-detection receiver. The goal of this paper is to obtain insights into the capacity and the structure of the capacity-achieving distribution for the channel under the amplitude constraint $\mathsf{A}$ and in the presence of dark current $λ$. Using recent theoretical progress on the structure of the capacity-achieving distribution, this paper develops a numerical algorithm, based on the gradient ascent and Blahut-Arimoto algorithms, for computing the capacity and the capacity-achieving distribution. The algorithm is used to perform extensive numerical simulations for various regimes of $\mathsf{A}$ and $λ$.

cs.IT↗

Scalar Gaussian Wiretap Channel with Peak Amplitude Constraint: Numerical Computation of the Optimal Input Distribution

This paper studies a scalar Gaussian wiretap channel where instead of an average input power constraint, we consider a peak amplitude constraint on the input. The goal is to obtain insights into the secrecy-capacity and the structure of the secrecy-capacity-achieving distribution. Capitalizing on the recent theoretical progress on the structure of the secrecy-capacity-achieving distribution, this paper develops a numerical procedure, based on the gradient ascent algorithm and a version of the Blahut-Arimoto algorithm, for computing the secrecy-capacity and the secrecy-capacity-achieving input and output distributions.

cs.IT↗

Scalar Gaussian Wiretap Channel: Properties of the Support Size of the Secrecy-Capacity-Achieving Distribution

This work studies the secrecy-capacity of a scalar-Gaussian wiretap channel with an amplitude constraint on the input. It is known that for this channel, the secrecy-capacity-achieving distribution is discrete with finitely many points. This work improves such result by showing an upper bound of the order $\frac{\mathsf{A}}{σ_1^2}$ where $\mathsf{A}$ is the amplitude constraint and $σ_1^2$ is the variance of the Gaussian noise over the legitimate channel.

cs.IT↗

A General Derivative Identity for the Conditional Expectation with Focus on the Exponential Family

Consider a pair of random vectors $(\mathbf{X},\mathbf{Y}) $ and the conditional expectation operator $\mathbb{E}[\mathbf{X}|\mathbf{Y}=\mathbf{y}]$. This work studies analytic properties of the conditional expectation by characterizing various derivative identities. The paper consists of two parts. In the first part of the paper, a general derivative identity for the conditional expectation is derived. Specifically, for the Markov chain $\mathbf{U} \leftrightarrow \mathbf{X} \leftrightarrow \mathbf{Y}$, a compact expression for the Jacobian matrix of $\mathbb{E}[\mathbf{U}|\mathbf{Y}=\mathbf{y}]$ is derived. In the second part of the paper, the main identity is specialized to the exponential family. Moreover, via various choices of the random vector $\mathbf{U}$, the new identity is used to recover and generalize several known identities and derive some new ones. As a first example, a connection between the Jacobian of $ \mathbb{E}[\mathbf{X}|\mathbf{Y}=\mathbf{y}]$ and the conditional variance is established. As a second example, a recursive expression between higher order conditional expectations is found, which is shown to lead to a generalization of the Tweedy's identity. Finally, as a third example, it is shown that the $k$-th order derivative of the conditional expectation is proportional to the $(k+1)$-th order conditional cumulant.

math.PR↗

Properties of the Support of the Capacity-Achieving Distribution of the Amplitude-Constrained Poisson Noise Channel

This work considers a Poisson noise channel with an amplitude constraint. It is well-known that the capacity-achieving input distribution for this channel is discrete with finitely many points. We sharpen this result by introducing upper and lower bounds on the number of mass points. Concretely, an upper bound of order $\mathsf{A} \log^2(\mathsf{A})$ and a lower bound of order $\sqrt{\mathsf{A}}$ are established where $\mathsf{A}$ is the constraint on the input amplitude. In addition, along the way, we show several other properties of the capacity and capacity-achieving distribution. For example, it is shown that the capacity is equal to $ - \log P_{Y^\star}(0)$ where $P_{Y^\star}$ is the optimal output distribution. Moreover, an upper bound on the values of the probability masses of the capacity-achieving distribution and a lower bound on the probability of the largest mass point are established. Furthermore, on the per-symbol basis, a nonvanishing lower bound on the probability of error for detecting the capacity-achieving distribution is established under the maximum a posteriori rule.

cs.IT↗

Consistent Density Estimation Under Discrete Mixture Models

This work considers a problem of estimating a mixing probability density $f$ in the setting of discrete mixture models. The paper consists of three parts. The first part focuses on the construction of an $L_1$ consistent estimator of $f$. In particular, under the assumptions that the probability measure $μ$ of the observation is atomic, and the map from $f$ to $μ$ is bijective, it is shown that there exists an estimator $f_n$ such that for every density $f$ $\lim_{n\to \infty} \mathbb{E} \left[ \int |f_n -f | \right]=0$. The second part discusses the implementation details. Specifically, it is shown that the consistency for every $f$ can be attained with a computationally feasible estimator. The third part, as a study case, considers a Poisson mixture model. In particular, it is shown that in the Poisson noise setting, the bijection condition holds and, hence, estimation can be performed consistently for every $f$.

cs.IT↗