SearcharxivSearch

arXiv · 2306.08800

Modules and PQ-trees in Robinson spaces

Abstract

A Robinson space is a dissimilarity space $(X,d)$ on $n$ points for which there exists a compatible order, {\it i.e.} a total order $<$ on $X$ such that $x<y<z$ implies that $d(x,y)\le d(x,z)$ and $d(y,z)\leq d(x,z)$. Recognizing if a dissimilarity space is Robinson has numerous applications in seriation and classification. A PQ-tree is a classical data structure introduced by Booth and Lueker to compactly represent a set of related permutations on a set $X$. In particular, the set of all compatible orders of a Robinson space are encoded by a PQ-tree. An mmodule is a subset $M$ of $X$ which is not distinguishable from the outside of $M$, {\it i.e.} the distances from any point of $X\setminus M$ to all points of $M$ are the same. Mmodules define the mmodule-tree of a dissimilarity space $(X,d)$. Given $p\in X$, a $p$-copoint is a maximal mmodule not containing $p$. The $p$-copoints form a partition of $X\setminus \{p\}$. There exist two algorithms recognizing Robinson spaces in optimal $O(n^2)$ time. One uses PQ-trees and one uses a copoint partition of $(X, d)$. In this paper, we establish correspondences between the PQ-trees and the mmodule-trees of Robinson spaces. More precisely, we show how to construct the mmodule-tree of a Robinson dissimilarity from its PQ-tree and how to construct the PQ-tree from the odule-tree. To establish this translation, additionally to the previous notions, we introduce the notions of $\delta$-graph $G_\delta$ of a Robinson space and of $\delta$-mmodules, the connected components of $G_\delta$. We also use the dendrogram of the subdominant ultrametric of $d$. All these results also lead to optimal $O(n^2)$ time algorithms for constructing the PQ-tree and the mmodule tree of Robinson spaces.

Explore related subjects

Keep this discovery

BibTeXRIS

Mikhael Carmona, Victor Chepoi, Guyslain Naves, Pascal Préa. 2023-06-15. Modules and PQ-trees in Robinson spaces. https://arxiv.org/abs/2306.08800

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An FPTAS for Two-Machine Open-Shop Scheduling with a Single Unavailability Interval

We consider the two-machine open-shop scheduling problem in which one machine is unavailable during a fixed interval. We study the resumable setting: an operation interrupted by the unavailability interval may resume, without penalty, when the machine becomes available. The objective is to minimize the makespan. Although the problem is NP-hard and several approximation algorithms are known, whether it admits a fully polynomial-time approximation scheme (FPTAS) has remained open for two decades. We resolve this question affirmatively by giving the first FPTAS, thereby strengthening the previously known polynomial-time approximation scheme (PTAS). As an intermediate result, we develop a new pseudo-polynomial dynamic program with seven state dimensions, improving on the ten-dimensional formulation in the literature.

cs.DM

Generalized Graph Search Trees

Graph search algorithms and their corresponding graph search trees are commonly used in algorithmic graph theory. In recent years, the recognition problem of these graph search trees has received significant attention. So far, the research has focused on two types of search trees: first-in trees that behave like BFS-trees and last-in trees that behave like DFS-trees. The search tree paradigms differ from each other by the parent a vertex is connected to. In first-in trees, it is the first visited neighbor, while in last-in trees it is the last neighbor visited before that vertex. Here, we will generalize these concepts of graph search trees by allowing every preceding neighbor of a vertex to be the parent. We study the complexity of the recognition problem of these generalized graph search trees. We present NP-completeness proofs for most searches. We also show that the problem is trivial for Generic Search and polynomial-time solvable for several searches on bipartite graphs and chordal graphs. We also study the question how fixing the start vertex influences the complexity of the problem.

cs.DM

The exact asymptotic constant in the metric dimension of Jaccard space

Let $X$ be a finite set with $|X|=n$ and let $\mathrm{Jac}(a,b)=|a\,\triangle\, b|/|a\cup b|$ be the Jaccard distance on the power set $2^X$. Lladser and Paradise recently proved that the metric dimension of $(2^X,\mathrm{Jac})$ is $\Theta(n/\ln n)$, with the constant left open; their bounds are $(\ln 2)\,n/\ln n\lesssim \beta(2^X,\mathrm{Jac})\lesssim 2\ln(2e)\,n/\ln n$. We determine the constant: \[ \beta(2^X,\mathrm{Jac})=\frac{2n}{\log_2 n}\,(1+o(1))=(2\ln 2)\,\frac{n}{\ln n}\,(1+o(1)). \] The proof identifies the problem, on each ``slice'' of subsets of fixed cardinality, with the Erd\H{o}s--R\'enyi coin-weighing problem for a spring scale (the problem of \emph{detecting matrices}). The lower bound is the Erd\H{o}s--R\'enyi entropy argument applied to the middle slice; the upper bound follows from the explicit detecting families of Lindstr\"om and of Cantor and Mills, augmented by a single extra landmark that reveals cardinality.

cs.DM