SearcharxivSearch

arXiv subjects

Alexander Tolmachev

Publications and source records attributed to Alexander Tolmachev.

10 recordsLinked to original sources

Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation

Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate estimators on simplified, low-dimensional distributions, leaving their performance on complex, realistic data largely unexplored. We address this gap with a comprehensive benchmarking framework grounded in a unified copula-theoretic perspective that subsumes existing benchmarks as special cases. Within this framework, we propose two complementary families of tests: a copula-first family that systematically varies ground-truth MI, dimensionality, and marginal complexity using synthetic and flow-based transformations; and a marginals-first family that couples real-world image data with controlled dependency structures, extending the classic same-class-pairing paradigm. We use this suite to extensively evaluate three classes of estimators: non-parametric, discriminative, and generative. Contrary to prevailing assumptions, our results indicate that there is no universal winner: each category can systematically outperform all other estimators under specific setups. By analyzing these cases, we identify fundamental estimation barriers and propose new tests that more effectively stress these specific limitations. We share the open source code at https://github.com/VanessB/mutinfo.

cs.LG

Reducing the upper bound for the Borsuk number in $\mathbb{R}^4$ to 8

The Borsuk number $b(n)$ of $n$-dimensional Euclidean space $\mathbb{R}^n$ is the smallest integer such that any set $F \subset \mathbb{R}^n$ of unit diameter can be partitioned into $b(n)$ subsets of strictly smaller diameter. For $n=4$, the best known upper bound $b(4) \leq 9$ follows from a construction by M. Lassak (1982). In the present paper, we construct partitions of several variants of the truncated Lassak cover into 8 parts of diameter less than 1, thereby showing that $b(4) \leq 8$.

math.MG

GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver

While diffusion models achieve state-of-the-art generation quality, they still suffer from computationally expensive sampling. Recent works address this issue with gradient-based optimization methods that distill a few-step ODE diffusion solver from the full sampling process, reducing the number of function evaluations from dozens to just a few. However, these approaches often rely on intricate training techniques and do not explicitly focus on preserving fine-grained details. In this paper, we introduce the Generalized Solver: a simple parameterization of the ODE sampler that does not require additional training tricks and improves quality over existing approaches. We further combine the original distillation loss with adversarial training, which mitigates artifacts and enhances detail fidelity. We call the resulting method the Generalized Adversarial Solver and demonstrate its superior performance compared to existing solver training methods under similar resource constraints. Code is available at https://github.com/3145tttt/GAS.

cs.CV

On lower bounds of the density of planar periodic sets without unit distances

Determining the maximal density $m_1(\mathbb{R}^2)$ of planar sets without unit distances is a fundamental problem in combinatorial geometry. This paper investigates lower bounds for this quantity. We introduce a novel approach to estimating $m_1(\mathbb{R}^2)$ by reformulating the problem as a Maximal Independent Set (MIS) problem on graphs constructed from flat torus, focusing on periodic sets with respect to two non-collinear vectors. Our experimental results, supported by theoretical justifications of proposed method, demonstrate that for a sufficiently wide range of parameters this approach does not improve the known lower bound $0.22936 \le m_1(\mathbb{R}^2)$. The best discrete sets found are approximations of Croft's construction. In addition, several open source software packages for MIS problem are compared on this task.

math.MG

Efficient Distribution Matching of Representations via Noise-Injected Deep InfoMax

Deep InfoMax (DIM) is a well-established method for self-supervised representation learning (SSRL) based on maximization of the mutual information between the input and the output of a deep neural network encoder. Despite the DIM and contrastive SSRL in general being well-explored, the task of learning representations conforming to a specific distribution (i.e., distribution matching, DM) is still under-addressed. Motivated by the importance of DM to several downstream tasks (including generative modeling, disentanglement, outliers detection and other), we enhance DIM to enable automatic matching of learned representations to a selected prior distribution. To achieve this, we propose injecting an independent noise into the normalized outputs of the encoder, while keeping the same InfoMax training objective. We show that such modification allows for learning uniformly and normally distributed representations, as well as representations of other absolutely continuous distributions. Our approach is tested on various downstream tasks. The results indicate a moderate trade-off between the performance on the downstream tasks and quality of DM.

cs.LG

Mutual Information Estimation via Normalizing Flows

We propose a novel approach to the problem of mutual information (MI) estimation via introducing a family of estimators based on normalizing flows. The estimator maps original data to the target distribution, for which MI is easier to estimate. We additionally explore the target distributions with known closed-form expressions for MI. Theoretical guarantees are provided to demonstrate that our approach yields MI estimates for the original data. Experiments with high-dimensional data are conducted to highlight the practical advantages of the proposed method.

cs.LG

Optimal partitions of the flat torus into parts of smaller diameter

We consider the problem of partitioning a two-dimensional flat torus $T^2$ into $m$ sets in order to minimize the maximal diameter of a part. For $m \leqslant 25$ we give numerical estimates for the maximal diameter $d_m(T^2)$ at which the partition exists. Several approaches are proposed to obtain such estimates. In particular, we use the search for mesh partitions via the SAT solver, the global optimization approach for polygonal partitions, and the optimization of periodic hexagonal tilings. For $m=3$, the exact estimate is proved using elementary topological reasoning.

math.MG

Information Bottleneck Analysis of Deep Neural Networks via Lossy Compression

The Information Bottleneck (IB) principle offers an information-theoretic framework for analyzing the training process of deep neural networks (DNNs). Its essence lies in tracking the dynamics of two mutual information (MI) values: between the hidden layer output and the DNN input/target. According to the hypothesis put forth by Shwartz-Ziv & Tishby (2017), the training process consists of two distinct phases: fitting and compression. The latter phase is believed to account for the good generalization performance exhibited by DNNs. Due to the challenging nature of estimating MI between high-dimensional random vectors, this hypothesis was only partially verified for NNs of tiny sizes or specific types, such as quantized NNs. In this paper, we introduce a framework for conducting IB analysis of general NNs. Our approach leverages the stochastic NN method proposed by Goldfeld et al. (2019) and incorporates a compression step to overcome the obstacles associated with high dimensionality. In other words, we estimate the MI between the compressed representations of high-dimensional random vectors. The proposed method is supported by both theoretical and practical justifications. Notably, we demonstrate the accuracy of our estimator through synthetic experiments featuring predefined MI values and comparison with MINE (Belghazi et al., 2018). Finally, we perform IB analysis on a close-to-real-scale convolutional DNN, which reveals new features of the MI dynamics.

cs.LG

Coverings of planar and three-dimensional sets with subsets of smaller diameter

Quantitative estimates related to the classical Borsuk problem of splitting set in Euclidean space into subsets of smaller diameter are considered. For a given $k$ there is a minimal diameter of subsets at which there exists a covering with $k$ subsets of any planar set of unit diameter. In order to find an upper estimate of the minimal diameter we propose an algorithm for finding sub-optimal partitions. In the cases $10 \leqslant k \leqslant 17$ some upper and lower estimates of the minimal diameter are improved. Another result is that any set $M \subset \mathbb{R}^3$ of a unit diameter can be partitioned into four subsets of a diameter not greater than $0.966$.

math.MG

H2O MegaMasers: a RadioAstron success story

The RadioAstron space-VLBI mission has successfully detected extragalactic H2O MegaMaser emission regions at very long Earth to space baselines ranging between 1.4 and 26.7 Earth Diameters (ED). The preliminary results for two galaxies, NGC3079 and NGC4258, at baselines longer than one ED indicate masering environments and excitation conditions in these galaxies that are distinctly different. Further observations of NGC4258 at longer baselines will reveal more of the physics of individual emission regions.

astro-ph.GA