SearcharxivSearch

arXiv subjects

Amy Glen

Publications and source records attributed to Amy Glen.

At least 19 recordsLinked to original sources

Structural Generalizability: The Case of Similarity Search

Graph similarity search algorithms usually leverage the structural properties of a database. Hence, these algorithms are effective only on some structural variations of the data and are ineffective on other forms, which makes them hard to use. Ideally, one would like to design a data analytics algorithm that is structurally robust, i.e., it returns essentially the same accurate results over all possible structural variations of a dataset. We propose a novel approach to create a structurally robust similarity search algorithm over graph databases. We leverage the classic insight in the database literature that schematic variations are caused by having constraints in the database. We then present RelSim algorithm which is provably structurally robust under these variations. Our empirical studies show that our proposed algorithms are structurally robust while being efficient and as effective as or more effective than the state-of-the-art similarity search algorithms.

cs.DB

More properties of the Fibonacci word on an infinite alphabet

Recently the Fibonacci word $W$ on an infinite alphabet was introduced by [Zhang et al., Electronic J. Combinatorics 24-2 (2017) #P2.52] as a fixed point of the morphism $ϕ: (2i) \mapsto (2i)(2i+ 1),\ (2i+ 1) \mapsto (2i+ 2)$ over all $i \in \mathbb{N}$. In this paper we investigate the occurrence of squares, palindromes, and Lyndon factors in this infinite word.

math.CO

Palindromes in starlike trees

In this note, we obtain an upper bound on the maximum number of distinct non-empty palindromes in starlike trees. This bound implies, in particular, that there are at most $4n$ distinct non-empty palindromes in a starlike tree with three branches each of length $n$. For such starlike trees labelled with a binary alphabet, we sharpen the upper bound to $4n-1$ and conjecture that the actual maximum is $4n-2$. It is intriguing that this simple conjecture seems difficult to prove, in contrast to the straightforward proof of the bound.

math.CO

Counting Lyndon factors

In this paper, we determine the maximum number of distinct Lyndon factors that a word of length $n$ can contain. We also derive formulas for the expected total number of Lyndon factors in a word of length $n$ on an alphabet of size $σ$, as well as the expected number of distinct Lyndon factors in such a word. The minimum number of distinct Lyndon factors in a word of length $n$ is $1$ and the minimum total number is $n$, with both bounds being achieved by $x^n$ where $x$ is a letter. A more interesting question to ask is what is the minimum number of distinct Lyndon factors in a Lyndon word of length $n$? In this direction, it is known (Saari, 2014) that an optimal lower bound for the number of distinct Lyndon factors in a Lyndon word of length $n$ is $\lceil\log_ϕ(n) + 1\rceil$, where $ϕ$ denotes the golden ratio $(1 + \sqrt{5})/2$. Moreover, this lower bound is attained by the so-called finite "Fibonacci Lyndon words", which are precisely the Lyndon factors of the well-known "infinite Fibonacci word" -- a special example of a "infinite Sturmian word". Saari (2014) conjectured that if $w$ is Lyndon word of length $n$, $n\ne 6$, containing the least number of distinct Lyndon factors over all Lyndon words of the same length, then $w$ is a Christoffel word (i.e., a Lyndon factor of an infinite Sturmian word). We give a counterexample to this conjecture. Furthermore, we generalise Saari's result on the number of distinct Lyndon factors of a Fibonacci Lyndon word by determining the number of distinct Lyndon factors of a given Christoffel word. We end with two open problems.

math.CO

Generalized trapezoidal words

The factor complexity function $C_w(n)$ of a finite or infinite word $w$ counts the number of distinct factors of $w$ of length $n$ for each $n \ge 0$. A finite word $w$ of length $|w|$ is said to be trapezoidal if the graph of its factor complexity $C_w(n)$ as a function of $n$ (for $0 \leq n \leq |w|$) is that of a regular trapezoid (or possibly an isosceles triangle); that is, $C_w(n)$ increases by 1 with each $n$ on some interval of length $r$, then $C_w(n)$ is constant on some interval of length $s$, and finally $C_w(n)$ decreases by 1 with each $n$ on an interval of the same length $r$. Necessarily $C_w(1)=2$ (since there is one factor of length $0$, namely the empty word), so any trapezoidal word is on a binary alphabet. Trapezoidal words were first introduced by de Luca (1999) when studying the behaviour of the factor complexity of finite Sturmian words, i.e., factors of infinite "cutting sequences", obtained by coding the sequence of cuts in an integer lattice over the positive quadrant of $\mathbb{R}^2$ made by a line of irrational slope. Every finite Sturmian word is trapezoidal, but not conversely. However, both families of words (trapezoidal and Sturmian) are special classes of so-called "rich words" (also known as "full words") - a wider family of finite and infinite words characterized by containing the maximal number of palindromes - studied in depth by the first author and others in 2009. In this paper, we introduce a natural generalization of trapezoidal words over an arbitrary finite alphabet $\mathcal{A}$, called generalized trapezoidal words (or GT-words for short). In particular, we study combinatorial and structural properties of this new class of words, and we show that, unlike the binary case, not all GT-words are rich in palindromes when $|\mathcal{A}| \geq 3$, but we can describe all those that are rich.

math.CO

The total run length of a word

A run in a word is a periodic factor whose length is at least twice its period and which cannot be extended to the left or right (by a letter) to a factor with greater period. In recent years a great deal of work has been done on estimating the maximum number of runs that can occur in a word of length $n$. A number of associated problems have also been investigated. In this paper we consider a new variation on the theme. We say that the total run length (TRL) of a word is the sum of the lengths of the runs in the word and that $τ(n)$ is the maximum TRL over all words of length $n$. We show that $n^2/8 < τ(n) < 47n^2/72 + 2n$ for all $n$. We also give a formula for the average total run length of words of length $n$ over an alphabet of size $α$, and some other results.

math.CO

Extremal properties of (epi)Sturmian sequences and distribution modulo 1

Starting from a study of Y. Bugeaud and A. Dubickas (2005) on a question in distribution of real numbers modulo 1 via combinatorics on words, we survey some combinatorial properties of (epi)Sturmian sequences and distribution modulo 1 in connection to their work. In particular we focus on extremal properties of (epi)Sturmian sequences, some of which have been rediscovered several times.

math.NT

Crucial words for abelian powers

A word is "crucial" with respect to a given set of "prohibited words" (or simply "prohibitions") if it avoids the prohibitions but it cannot be extended to the right by any letter of its alphabet without creating a prohibition. A "minimal crucial word" is a crucial word of the shortest length. A word W contains an "abelian k-th power" if W has a factor of the form X_1X_2...X_k where X_i is a permutation of X_1 for 2<= i <= k. When k=2 or 3, one deals with "abelian squares" and "abelian cubes", respectively. In 2004 (arXiv:math/0205217), Evdokimov and Kitaev showed that a minimal crucial word over an n-letter alphabet A_n = {1,2,..., n} avoiding abelian squares has length 4n-7 for n >= 3. In this paper we show that a minimal crucial word over A_n avoiding abelian cubes has length 9n-13 for n >= 5, and it has length 2, 5, 11, and 20 for n=1, 2, 3, and 4, respectively. Moreover, for n >= 4 and k >= 2, we give a construction of length k^2(n-1)-k-1 of a crucial word over A_n avoiding abelian k-th powers. This construction gives the minimal length for k=2 and k=3. For k >= 4 and n >= 5, we provide a lower bound for the length of crucial words over A_n avoiding abelian k-th powers.

math.CO

Quasiperiodic and Lyndon episturmian words

Recently the second two authors characterized quasiperiodic Sturmian words, proving that a Sturmian word is non-quasiperiodic if and only if it is an infinite Lyndon word. Here we extend this study to episturmian words (a natural generalization of Sturmian words) by describing all the quasiperiods of an episturmian word, which yields a characterization of quasiperiodic episturmian words in terms of their "directive words". Even further, we establish a complete characterization of all episturmian words that are Lyndon words. Our main results show that, unlike the Sturmian case, there is a much wider class of episturmian words that are non-quasiperiodic, besides those that are infinite Lyndon words. Our key tools are morphisms and directive words, in particular "normalized" directive words, which we introduced in an earlier paper. Also of importance is the use of "return words" to characterize quasiperiodic episturmian words, since such a method could be useful in other contexts.

math.CO

Episturmian words: a survey

In this paper, we survey the rich theory of infinite episturmian words which generalize to any finite alphabet, in a rather resembling way, the well-known family of Sturmian words on two letters. After recalling definitions and basic properties, we consider episturmian morphisms that allow for a deeper study of these words. Some properties of factors are described, including factor complexity, palindromes, fractional powers, frequencies, and return words. We also consider lexicographical properties of episturmian words, as well as their connection to the balance property, and related notions such as finite episturmian words, Arnoux-Rauzy sequences, and "episkew words" that generalize the skew words of Morse and Hedlund.

math.CO

A new characteristic property of rich words

Originally introduced and studied by the third and fourth authors together with J. Justin and S. Widmer in arXiv:0801.1656, rich words constitute a new class of finite and infinite words characterized by containing the maximal number of distinct palindromes. Several characterizations of rich words have already been established. A particularly nice characteristic property is that all 'complete returns' to palindromes are palindromes. In this note, we prove that rich words are also characterized by the property that each factor is uniquely determined by its longest palindromic prefix and its longest palindromic suffix.

math.CO

A connection between palindromic and factor complexity using return words

In this paper we prove that for any infinite word W whose set of factors is closed under reversal, the following conditions are equivalent: (I) all complete returns to palindromes are palindromes; (II) P(n) + P(n+1) = C(n+1) - C(n) + 2 for all n, where P (resp. C) denotes the palindromic complexity (resp. factor complexity) function of W, which counts the number of distinct palindromic factors (resp. factors) of each length in W.

math.CO

Palindromic Richness

In this paper, we study combinatorial and structural properties of a new class of finite and infinite words that are 'rich' in palindromes in the utmost sense. A characteristic property of so-called "rich words" is that all complete returns to any palindromic factor are themselves palindromes. These words encompass the well-known episturmian words, originally introduced by the second author together with X. Droubay and G. Pirillo in 2001. Other examples of rich words have appeared in many different contexts. Here we present the first unified approach to the study of this intriguing family of words. Amongst our main results, we give an explicit description of the periodic rich infinite words and show that the recurrent balanced rich infinite words coincide with the balanced episturmian words. We also consider two wider classes of infinite words, namely "weakly rich words" and almost rich words (both strictly contain all rich words, but neither one is contained in the other). In particular, we classify all recurrent balanced weakly rich words. As a consequence, we show that any such word on at least three letters is necessarily episturmian; hence weakly rich words obey Fraenkel's conjecture. Likewise, we prove that a certain class of almost rich words obeys Fraenkel's conjecture by showing that the recurrent balanced ones are episturmian or contain at least two distinct letters with the same frequency. Lastly, we study the action of morphisms on (almost) rich words with particular interest in morphisms that preserve (almost) richness. Such morphisms belong to the class of "P-morphisms" that was introduced by A. Hof, O. Knill, and B. Simon in 1995.

math.CO

Rich, Sturmian, and trapezoidal words

In this paper we explore various interconnections between rich words, Sturmian words, and trapezoidal words. Rich words, first introduced in arXiv:0801.1656 by the second and third authors together with J. Justin and S. Widmer, constitute a new class of finite and infinite words characterized by having the maximal number of palindromic factors. Every finite Sturmian word is rich, but not conversely. Trapezoidal words were first introduced by the first author in studying the behavior of the subword complexity of finite Sturmian words. Unfortunately this property does not characterize finite Sturmian words. In this note we show that the only trapezoidal palindromes are Sturmian. More generally we show that Sturmian palindromes can be characterized either in terms of their subword complexity (the trapezoidal property) or in terms of their palindromic complexity. We also obtain a similar characterization of rich palindromes in terms of a relation between palindromic complexity and subword complexity.

math.CO

Directive words of episturmian words: equivalences and normalization

Episturmian morphisms constitute a powerful tool to study episturmian words. Indeed, any episturmian word can be infinitely decomposed over the set of pure episturmian morphisms. Thus, an episturmian word can be defined by one of its morphic decompositions or, equivalently, by a certain directive word. Here we characterize pairs of words directing a common episturmian word. We also propose a way to uniquely define any episturmian word through a normalization of its directive words. As a consequence of these results, we characterize episturmian words having a unique directive word.

cs.DM

On the critical exponent of generalized Thue-Morse words

For certain generalized Thue-Morse words t, we compute the "critical exponent", i.e., the supremum of the set of rational numbers that are exponents of powers in t, and determine exactly the occurrences of powers realizing it.

math.CO

Conjugates of characteristic Sturmian words generated by morphisms

This article is concerned with characteristic Sturmian words of slope $α$ and $1-α$ (denoted by $c_α$ and $c_{1-α}$ respectively), where $α\in (0,1)$ is an irrational number such that $α= [0;1+d_1,\bar{d_2,...,d_n}]$ with $d_n \geq d_1 \geq 1$. It is known that both $c_α$ and $c_{1-α}$ are fixed points of non-trivial (standard) morphisms $σ$ and $\hatσ$, respectively, if and only if $α$ has a continued fraction expansion as above. Accordingly, such words $c_α$ and $c_{1-α}$ are generated by the respective morphisms $σ$ and $\hatσ$. For the particular case when $α= [0;2,\bar{r}]$ ($r\geq1$), we give a decomposition of each conjugate of $c_α$ (and hence $c_{1-α}$) into generalized adjoining singular words, by considering conjugates of powers of the standard morphism $σ$ by which it is generated. This extends a recent result of Levé and S\ee bold on conjugates of the infinite Fibonacci word.

math.CO