Searcharxiv⌕ Search

arXiv subjects

Daniel Stefankovic

Publications and source records attributed to Daniel Stefankovic.

38 records · Page 3Linked to original sources

Adaptive Simulated Annealing: A Near-optimal Connection between Sampling and Counting

We present a near-optimal reduction from approximately counting the cardinality of a discrete set to approximately sampling elements of the set. An important application of our work is to approximating the partition function $Z$ of a discrete system, such as the Ising model, matchings or colorings of a graph. The typical approach to estimating the partition function $Z(β^*)$ at some desired inverse temperature $β^*$ is to define a sequence, which we call a {\em cooling schedule}, $β_0=0<β_1<...<β_\ell=β^*$ where Z(0) is trivial to compute and the ratios $Z(β_{i+1})/Z(β_i)$ are easy to estimate by sampling from the distribution corresponding to $Z(β_i)$. Previous approaches required a cooling schedule of length $O^*(\ln{A})$ where $A=Z(0)$, thereby ensuring that each ratio $Z(β_{i+1})/Z(β_i)$ is bounded. We present a cooling schedule of length $\ell=O^*(\sqrt{\ln{A}})$. For well-studied problems such as estimating the partition function of the Ising model, or approximating the number of colorings or matchings of a graph, our cooling schedule is of length $O^*(\sqrt{n})$, which implies an overall savings of $O^*(n)$ in the running time of the approximate counting algorithm (since roughly $\ell$ samples are needed to estimate each ratio).

cs.DS↗

Phylogeny of Mixture Models: Robustness of Maximum Likelihood and Non-identifiable Distributions

We address phylogenetic reconstruction when the data is generated from a mixture distribution. Such topics have gained considerable attention in the biological community with the clear evidence of heterogeneity of mutation rates. In our work, we consider data coming from a mixture of trees which share a common topology, but differ in their edge weights (i.e., branch lengths). We first show the pitfalls of popular methods, including maximum likelihood and Markov chain Monte Carlo algorithms. We then determine in which evolutionary models, reconstructing the tree topology, under a mixture distribution, is (im)possible. We prove that every model whose transition matrices can be parameterized by an open set of multi-linear polynomials, either has non-identifiable mixture distributions, in which case reconstruction is impossible in general, or there exist linear tests which identify the topology. This duality theorem, relies on our notion of linear tests and uses ideas from convex programming duality. Linear tests are closely related to linear invariants, which were first introduced by Lake, and are natural from an algebraic geometry perspective.

q-bio.PE↗