Generating and generalizing MSTD sets through Markov processes
The classical More Sums Than Differences (MSTD) problem studies finite sets $A\subset\{0,1,\ldots,n\}$ for which $|A+A|>|A-A|$, where $A+A=\{a_1+a_2:a_1,a_2\in A\}$ and $A-A=\{a_1-a_2:a_1,a_2\in A\}$. As addition is commutative and subtraction is not, it was conjectured that as $n\to\infty$ almost all subsets $A$ chosen uniformly from the power set of $\{0,1,\ldots,n\}$ are difference-dominated, and it was thus a surprise when Martin and O'Bryant proved a positive percentage of sets are sum-dominant. We greatly generalize this model by introducing a Markov-chain framework, where the classical MSTD model is now just a special case. Let $(X_i)_{i=0}^n$ be a stationary two-state Markov chain on $\{0,1\}$ with transition probabilities $P(0,0)=p$ and $P(1,1)=q$, where $p,q\in(0,1)$. We include $i$ in $A$ exactly when $X_i=1$, and define $A=\{i\in\{0,\ldots,n\}:X_i=1\}$. The usual independent Bernoulli model is recovered when consecutive inclusion decisions are independent, equivalently when $p=1-q$. In particular, the uniformly random subset model corresponds to $p=q=1/2$. Using the fringe-middle method from the MSTD literature, we show that the middle sums and differences are filled with high probability, so the comparison between $|A+A|$ and $|A-A|$ is again governed by endpoint fringes. By fringe manipulation, we prove that the probabilities of sum-dominant, difference-dominant, and balanced sets tend to strictly positive limits as $n\to\infty$. We also give numerical estimates of these three probabilities for finite $n$ over a range of values of $p$ and $q$. Through combinatorial methods, we find a closed-form expression for $\mathbb{E}[|A-A|-|A+A|]$ as $n\to\infty$.