SearcharxivSearch

arXiv subjects

Marco Dalai

Publications and source records attributed to Marco Dalai.

At least 19 recordsLinked to original sources

The Expurgated Error Exponent is Not Universally Achievable

We study the universal attainability of the expurgated error exponent for discrete memoryless channels (DMCs). While the random-coding exponent is known to be universally attainable via maximum mutual information (MMI) decoding for DMCs, it remains open whether the expurgated exponent can be attained universally. We show that this is not the case in general. Specifically, we construct a family of DMCs for which no single sequence of codes can attain the expurgated exponent simultaneously for all channels in the family, even at rate zero. In addition, for the same channel family, we show that MMI decoding fails to achieve the expurgated exponent for any channel in the family.

cs.IT

Zero-Error List Decoding for Classical-Quantum Channels

The aim of this work is to study the zero-error capacity of pure-state classical-quantum channels in the setting of list decoding. We provide an achievability bound for list-size two and a converse bound holding for every fixed list size. The two bounds coincide for channels whose pairwise absolute state overlaps form a positive semi-definite matrix. Finally, we discuss a remarkable peculiarity of the classical-quantum case: differently from the fully classical setting, the rate at which the sphere-packing bound diverges might not be achievable by zero-error list codes, even when we take the limit of fixed but arbitrarily large list size.

quant-ph

Bounds on $k$-hash distances and rates of linear codes

In this paper, we bound the rate of linear codes in $\mathbb{F}_q^n$ with the property that any $k\leq q$ codewords are all simultaneously distinct in at least $d_k$ coordinates. For the case of particular interest $q=k=3$ we recover, with a simpler proof, state of the art results in the case $d_3=1$ and new bounds for $d_3>1$. We finally discuss some related open problems on the list-decoding zero-error capacity of discrete memoryless channels.

cs.IT

End-to-End Semantic Preservation in Text-Aware Image Compression Systems

Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. We further extend this study to general-purpose encoders, exploring their capacity to preserve hidden semantics under extreme compression. Instead of optimizing for visual fidelity, we examine whether compact, visually degraded representations can retain recoverable meaning through learned enhancement and recognition modules. Results demonstrate that semantic information can persist despite severe compression, bridging text-oriented compression and general-purpose semantic preservation in machine-centered image coding.

eess.IV

Definability of some $k$-ary Relations Over Second Order kinds of Logics

We consider the exprissibility in monadic second order logic of certain relations of importance in computer science. For integers $n\geq 1$ and $k\leq b$, a $k$-tuple of sequences in $\{0,1,\ldots, b-1\}^n$ are said to be $k$-hashed if there is a coordinate where they all differ. A set $\mathcal{C}$ of sequences is said to be a $k$-hash code if any $k$ distinct elements are $k$-hashed. Testing whether a code is $k$-hashing and determining the largest size of $k$-hash codes is an important problem in computer science. The use of general purpose solvers for this problem leads to question what minimal logic is needed to represent the problem. In this paper, we prove that the $k$-hashing relation on $k$-tuples is not definable in Monadic Second Order Logic (MSO), highlighting its limitations for this problem. Instead, the property can be expressed in extensions of the MSO that add the equi-cardinality relation.

math.LO

An Efficient Algorithm for Group Testing with Runlength Constraints

In this paper, we provide an efficient algorithm to construct almost optimal $(k,n,d)$-superimposed codes with runlength constraints. A $(k,n,d)$-superimposed code of length $t$ is a $t \times n$ binary matrix such that any two 1's in each column are separated by a run of at least $d$ 0's, and such that for any column $\mathbf{c}$ and any other $k-1$ columns, there exists a row where $\mathbf{c}$ has $1$ and all the remaining $k-1$ columns have $0$. These combinatorial structures were introduced by Agarwal et al. [1], in the context of Non-Adaptive Group Testing algorithms with runlength constraints. By using Moser and Tardos' constructive version of the Lovász Local Lemma, we provide an efficient randomized Las Vegas algorithm of complexity $Θ(t n^2)$ for the construction of $(k,n,d)$-superimposed codes of length $t=O(dk\log n +k^2\log n)$. We also show that the length of our codes is shorter, for $n$ sufficiently large, than that of the codes whose existence was proved in [1].

cs.IT

Upper bounds on the rate of linear $q$-ary $k$-hash codes

This paper presents new upper bounds on the rate of linear $k$-hash codes in $\mathbb{F}_q^n$, $q\geq k$, that is, codes with the property that any $k$ distinct codewords are all simultaneously distinct in at least one coordinate.

cs.IT

Bounds and Algorithms for Frameproof Codes and Related Combinatorial Structures

In this paper, we study upper bounds on the minimum length of frameproof codes introduced by Boneh and Shaw to protect copyrighted materials. A $q$-ary $(k,n)$-frameproof code of length $t$ is a $t \times n$ matrix having entries in $\{0,1,\ldots, q-1\}$ and with the property that for any column $\mathbf{c}$ and any other $k$ columns, there exists a row where the symbols of the $k$ columns are all different from the corresponding symbol (in the same row) of the column $\mathbf{c}$. In this paper, we show the existence of $q$-ary $(k,n)$-frameproof codes of length $t = O(\frac{k^2}{q} \log n)$ for $q \leq k$, using the Lovász Local Lemma, and of length $t = O(\frac{k}{\log(q/k)}\log(n/k))$ for $q > k$ using the expurgation method. Remarkably, for the practical case of $q \leq k$ our findings give codes whose length almost matches the lower bound $Ω(\frac{k^2}{q\log k} \log n)$ on the length of any $q$-ary $(k,n)$-frameproof code and, more importantly, allow us to derive an algorithm of complexity $O(t n^2)$ for the construction of such codes.

cs.IT

Variations on the Erdős distinct-sums problem

Let $\{a_1, . . . , a_n\}$ be a set of positive integers with $a_1 < \dots < a_n$ such that all $2^n$ subset sums are distinct. A famous conjecture by Erdős states that $a_n>c\cdot 2^n$ for some constant $c$, while the best result known to date is of the form $a_n>c\cdot 2^n/\sqrt{n}$. In this paper, we weaken the condition by requiring that only sums corresponding to subsets of size smaller than or equal to $λn$ be distinct. For this case, we derive lower and upper bounds on the smallest possible value of $a_n$.

math.CO

Achievable Rates and Algorithms for Group Testing with Runlength Constraints

In this paper, we study bounds on the minimum length of $(k,n,d)$-superimposed codes introduced by Agarwal et al. [1], in the context of Non-Adaptive Group Testing algorithms with runlength constraints. A $(k,n,d)$-superimposed code of length $t$ is a $t \times n$ binary matrix such that any two 1's in each column are separated by a run of at least $d$ 0's, and such that for any column $\mathbf{c}$ and any other $k-1$ columns, there exists a row where $\mathbf{c}$ has $1$ and all the remaining $k-1$ columns have $0$. Agarwal et al. proved the existence of such codes with $t=Θ(dk\log(n/k)+k^2\log(n/k))$. Here we investigate more in detail the coefficients in front of these two main terms as well as the role of lower order terms. We show that improvements can be obtained over the construction in [1] by using different constructions and by an appropriate exploitation of the Lovász Local Lemma in this context. Our findings also suggest $O(n^k)$ randomized Las Vegas algorithms for the construction of such codes. We also extend our results to Two-Stage Group Testing algorithms with runlength constraints.

cs.IT

Improved Bounds for $(b,k)$-hashing

For fixed integers $b\geq k$, a problem of relevant interest in computer science and combinatorics is that of determining the asymptotic growth, with $n$, of the largest set for which a $(b, k)$-hash family of $n$ functions exists. Equivalently, determining the asymptotic growth of a largest subset of $\{1,2,\ldots,b\}^n$ such that, for any $k$ distinct elements in the set, there is a coordinate where they all differ. An important asymptotic upper bound for general $b, k$, was derived by Fredman and Komlós in the '80s and improved for certain $b\neq k$ by Körner and Marton and by Arikan. Only very recently better bounds were derived for the general $b,k$ case by Guruswami and Riazanov while stronger results for small values of $b=k$ were obtained by Arikan, by Dalai, Guruswami and Radhakrishnan and by Costa and Dalai. In this paper, we both show how some of the latter results extend to $b\neq k$ and further strengthen the bounds for some specific small values of $b$ and $k$. The method we use, which depends on the reduction of an optimization problem to a finite number of cases, shows that further results might be obtained by refined arguments at the expense of higher complexity which could be reduced by using more sophisticated and optimized algorithmic approaches.

math.CO

Mismatched decoding reliability function at zero rate

We derive an upper bound on the reliability function of mismatched decoding for zero-rate codes. The bound is based on a result by Komlós that shows the existence of a subcode with certain symmetry properties. The bound is shown to coincide with the expurgated exponent at rate zero for a broad family of channel-decoding metric pairs.

cs.IT

A Revisitation of Low-Rate Bounds on the Reliability Function of Discrete Memoryless Channels for List Decoding

We revise the proof of low-rate upper bounds on the reliability function of discrete memoryless channels for ordinary and list-decoding schemes, in particular Berlekamp and Blinovsky's zero-rate bound, as well as Blahut's bound for low rates. The available proofs of the zero-rate bound devised by Berlekamp and Blinovsky are somehow complicated in that they contain in one form or another some cumbersome "non-standard" procedures or computations. Here we follow Blinovsky's idea of using a Ramsey-theoretic result by Komlos, and we complement it with some missing steps to present a proof which is rigorous and easier to inspect. Furthermore, we show how these techniques can be used to fix an error that invalidated the proof of Blahut's low-rate bound, which is here presented in an extended form for list decoding and for general channels.

cs.IT

A note on $\overline{2}$-separable codes and $B_2$ codes

We derive a simple proof, based on information theoretic inequalities, of an upper bound on the largest rates of $q$-ary $\overline{2}$-separable codes that improves recent results of Wang for any $q\geq 13$. For the case $q=2$, we recover a result of Lindström, but with a much simpler derivation. The method easily extends to give bounds on $B_2$ codes which, although not improving on Wang's results, use much simpler tools and might be useful for future applications.

math.CO

New upper bounds for $(b,k)$-hashing

For fixed integers $b\geq k$, the problem of perfect $(b,k)$-hashing asks for the asymptotic growth of largest subsets of $\{1,2,\ldots,b\}^n$ such that for any $k$ distinct elements in the set, there is a coordinate where they all differ. An important asymptotic upper bound for general $b, k$, was derived by Fredman and Komlós in the '80s and improved for certain $b\neq k$ by Körner and Marton and by Arikan. Only very recently better bounds were derived for the general $b,k$ case by Guruswami and Riazanov, while stronger results for small values of $b=k$ were obtained by Arikan, by Dalai, Guruswami and Radhakrishnan and by Costa and Dalai. In this paper, we both show how some of the latter results extend to $b\neq k$ and further strengthen the bounds for some specific small values of $b$ and $k$. The method we use, which depends on the reduction of an optimization problem to a finite number of cases, shows that further results might be obtained by refined arguments at the expense of higher complexity.

cs.IT

New bounds for perfect $k$-hashing

Let $C\subseteq \{1,\ldots,k\}^n$ be such that for any $k$ distinct elements of $C$ there exists a coordinate where they all differ simultaneously. Fredman and Komlós studied upper and lower bounds on the largest cardinality of such a set $C$, in particular proving that as $n\to\infty$, $|C|\leq \exp(n k!/k^{k-1}+o(n))$. Improvements over this result where first derived by different authors for $k=4$. More recently, Guruswami and Riazanov showed that the coefficient $k!/k^{k-1}$ is certainly not tight for any $k>3$, although they could only determine explicit improvements for $k=5,6$. For larger $k$, their method gives numerical values modulo a conjecture on the maxima of certain polynomials. In this paper, we first prove their conjecture, completing the explicit computation of an improvement over the Fredman-Komlós bound for any $k$. Then, we develop a different method which gives substantial improvements for $k=5,6$.

math.CO

A tour problem on a toroidal board

In this paper we study a tour problem that we came cross while studying biembeddings and Heffter arrays, see [D.S. Archdeacon, Heffter arrays and biembedding graphs on surfaces, Electron. J. Combin. 22 (2015) #P1.74]. Let $A$ be an $n\times m$ toroidal array consisting of filled cells and empty cells. Assume that an orientation $R=(r_1,\dots,r_n)$ of each row and $C=(c_1,\dots,c_m)$ of each column of $A$ is fixed. Given an initial filled cell $(i_1,j_1)$ consider the list $ L_{R,C}=((i_1,j_1),(i_2,j_2),\ldots,(i_k,j_k),$ $(i_{k+1},j_{k+1}),\ldots)$ where $j_{k+1}$ is the column index of the filled cell $(i_k,j_{k+1})$ of the row $R_{i_k}$ next to $(i_k,j_k)$ in the orientation $r_{i_k}$, and where $i_{k+1}$ is the row index of the filled cell of the column $C_{j_{k+1}}$ next to $(i_k,j_{k+1})$ in the orientation $c_{j_{k+1}}$. We propose the following "Crazy Knight's Tour Problem": Do there exist $R$ and $C$ such that the list $L_{R,C}$ covers all the filled cells of $A$? Here we provide a complete solution for the case with no empty cells and we obtain partial results for square arrays where the filled cells follow some specific regular patterns.

math.CO

A gap in the slice rank of $k$-tensors

The slice-rank method, introduced by Tao as a symmetrized version of the polynomial method of Croot, Lev and Pach and Ellenberg and Gijswijt, has proved to be a useful tool in a variety of combinatorial problems. Explicit tensors have been introduced in different contexts but little is known about the limitations of the method. In this paper, building upon a method presented by Tao and Sawin, it is proved that the asymptotic slice rank of any $k$-tensor in any field is either $1$ or at least $k/(k-1)^{(k-1)/k}$. This provides evidence that straight-forward application of the method cannot give useful results in certain problems for which non-trivial exponential bounds are already known. An example, actually a motivation for starting this work, is the problem of bounding the size of trifferent sets of sequences, which constitutes a long-standing open problem in information theory and in theoretical computer science.

math.CO