arXiv · 1911.02889
Towards Better Compressed Representations
Abstract
We introduce the problem of computing a parsing where each phrase is of length at most $m$ and which minimizes the zeroth order entropy of parsing. Based on the recent theoretical results we devise a heuristic for this problem. The solution has straightforward application in succinct text representations and gives practical improvements. Moreover the proposed heuristic yields structure whose size can be bounded both by $|S|H_{m-1}(S)$ and by $|S|/m(H_0(S) + \cdots + H_{m-1})$, where $H_{k}(S)$ is the $k$-th order empirical entropy of $S$. We also consider a similar problem in which the first-order entropy is minimized.
Explore related subjects
Keep this discovery
Michał Gańczorz. 2019-11-07. Towards Better Compressed Representations. https://arxiv.org/abs/1911.02889
Cite the original work for its findings. Save a collection to share your selection of sources.