SearcharxivSearch

arXiv subjects

Vera Koponen

Publications and source records attributed to Vera Koponen.

At least 19 recordsLinked to original sources

Exponential random graph models with soft clique constraints

Let $r\geq3$ be fixed, and let $\mathbf{G}_n$ be the set of all simple graphs with vertex set $[n]=\{1,\ldots,n\}$. We consider an exponential random graph model which gives higher probability to $G \in \mathbf{G}_n$ than to $H \in \mathbf{G}_n$ if $G$ has fewer $r$-cliques than $H$. But all graphs in $\mathbf{G}_n$ have positive probability. The degree to which graphs with fewer $r$-cliques are given higher probability is determined by a positive weight $w$. We prove that, asymptotically almost surely as $n \to \infty$, a random graph from $\mathbf{G}_n$ has a vertex partition into $r-1$ parts of roughly equal size, the density of edges between the parts is close to $1/2$, and for every $\varepsilon > 0$ the density of edges within any part is less than $\varepsilon$. The asymptotic structural properties are independent of the weight $w$ as long as it is positive. We also extend the result to the context of several clique sizes, each one with its own weight.

math.CO

A concentration result for multilayer feedforward neural networks

We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$ layers, $n$ neurons in the first layer (the input layer) and only one neuron, the output neuron, in the last layer. Very roughly formulated, the main result is that if the distribution of weights of connections from a layer to the next are, for all large $n$, approximated well by a fixed continuous (but otherwise arbitrary) curve which does not depend on $n$, and if the values of the $n$ input neurons are independently and identically distributed with a continuous probability density function, then there is a number $\psi$ such that for all $\varepsilon > 0$ the probability that the value of the output neuron is in $[\psi - \varepsilon, \psi + \varepsilon]$ tends to 1 as $n$ tends to infinity.

cs.AI

Random coloured digraphs defined by a Markov logic network

A Markov Logic Network (MLN) is a probabilistic relational model used in Statistical Relational Artificial Intelligence for defining a probability distribution on the set of possible worlds with domain $D$ for an arbitrary finite domain $D$. An MLN consists of soft constraints with associated weights which are nonnegative real numbers. In this study we consider a language speaking about a property $P(x)$ and a relation $R(x, y)$. We consider an MLN for which every Boolean combination of $P(x)$ and $R(x, y)$ is a soft constraint (with associated weight). Let $n$ denote the size (cardinality) of the domain. We show that, for every choice of weights, if the weights are scaled by $1/n$ then, for every first-order sentence $\varphi$, the probability that $\varphi$ holds tends to either 0 or 1 as $n \to \infty$; that is, a 0-1 law for first-order logic holds. Morover, the limit probability does {\em not} depend on the weights. If we instead use the standard semantics of MLNs, in the case of which the weights are {\em not} scaled, then the limit behaviour is more complicated and {\em depends} on the weights. With unscaled weights we get 7 qualitatively different cases which depend on the weights. In some cases we have a 0-1 law for first-order logic, in some cases not, but we may still have a convergence law. The influence of the weights on the asymptotic probability of a first-order sentence may be in the form of a sudden ``phase transition'' from one of the 7 cases to another. The presence of a convergence law has positive implications for inference on large domains.

math.LO

Notions of rank and independence in countably categorical theories

For an $\omega$-categorical theory $T$ and model $\mathcal{M}$ of $T$ we define a hierarchy of ranks, the $n$-ranks for $n < \omega$ which only care about imaginary elements ``up to level $n$'', where level $n$ contains every element of $M$ and every imaginary element that is an equivalence class of an $\emptyset$-definable equivalence relation on $n$-tuples of elements from $M$. Using the $n$-rank we define the notion of $n$-independence. For all $n < \omega$, the $n$-independence relation restricted to $M_n$ has all properties of an independence relation according to Kim and Pillay with the {\em possible exception} of the symmetry property. We prove that, given any $n < \omega$, if $\mathcal{M} \models T$ and the algebraic closure in $\mathcal{M}^{\mathrm{eq}}$ restricted to imaginary elements ``up to level $n$'' which have $n$-rank 1 (over some set of parameters) satisfies the exchange property, then $n$-independence is symmetric and hence an independence relation when restricted to $M_n$. Then we show that if $n$-independence is symmetric for all $n < \omega$, then $T$ is rosy. An application of this is that if $T$ has weak elimination of imaginaries and the algebraic closure in $\mathcal{M}$ restricted to elements of $M$ of 0-rank 1 (over some set of parameters from $M^{\mathrm{eq}}$) satisfies the exchange property, then $T$ is superrosy with finite U-thorn-rank.

math.LO

Domain size asymptotics for Markov logic networks

A Markov logic network (MLN) $\mathbb{M}$ determines a probability distribution $\mathbb{P}_n^\mathbb{M}$ on the set $\mathbf{W}_n$ of structures, or ``possible worlds'', with domain $\{1, \ldots, n\}$. We study the properties of such distributions as $n$ tends to infinity. We show that with mild assumptions on an MLN $\mathbb{M}$ with one soft constraint with an arbitrary positive weight the distribution $\mathbb{P}_n^\mathbb{M}$ will behave quite differently from the uniform distribution $\mathbb{P}_n^{uni}$ on $\mathbf{W}_n$ for all sufficiently large $n$. For a language with only one relation symbol $R$ which has arity 1 we give an almost complete characterization of the possible asymptotic behaviours of $\mathbb{P}_n^\mathbb{M}$ as $n \to \infty$, where $\mathbb{M}$ may be any MLN for this language. The asymptotic behaviour depends on the soft constraints and weights of the MLN. This characterization is used to show that if the language under consideration contains at least one relation symbol of arity 1 then the following holds: (a) There is an MLN $\mathbb{M}$ such that for every lifted Bayesian network (LBN) $\mathbb{G}$ there are infinitely many $n$ such that $\mathbb{M}$ and $\mathbb{G}$ determine different distributions on $\mathbf{W}_n$. (b) There is an LBN $\mathbb{G}$ such that for every MLN $\mathbb{M}$ there are infinitely many $n$ such that $\mathbb{G}$ and $\mathbb{M}$ determine different distributions on $\mathbf{W}_n$. We also show that, in the limit, the weight dimension and the domain size dimension may behave completely differently.

cs.AI

A convergence law for continuous logic and continuous structures with finite domains

We consider continuous relational structures with finite domain $[n] := \{1, \ldots, n\}$ and a many valued logic, $CLA$, with values in the unit interval and which uses continuous connectives and continuous aggregation functions. $CLA$ subsumes first-order logic on ``conventional'' finite structures. To each relation symbol $R$ and identity constraint $ic$ on a tuple the length of which matches the arity of $R$ we associate a continuous probability density function $\mu_R^{ic} : [0, 1] \to [0, \infty)$. We also consider a probability distribution on the set $\mathbf{W}_n$ of continuous structures with domain $[n]$ which is such that for every relation symbol $R$, identity constraint $ic$, and tuple $\bar{a}$ satisfying $ic$, the distribution of the value of $R(\bar{a})$ is given by $\mu_R^{ic}$, independently of the values for other relation symbols or other tuples. In this setting we prove that every formula in $CLA$ is asymptotically equivalent to a formula without any aggregation function. This is used to prove a convergence law for $CLA$ which reads as follows for formulas without free variables: If $\varphi \in CLA$ has no free variable and $I \subseteq [0, 1]$ is an interval, then there is $\alpha \in [0, 1]$ such that, as $n$ tends to infinity, the probability that the value of $\varphi$ is in $I$ tends to $\alpha$.

cs.LO

Random expansions of trees with bounded height

We consider a sequence $\mathbf{T} = (\mathcal{T}_n : n \in \mathbb{N}^+)$ of trees $\mathcal{T}_n$ where, for some $\Delta \in \mathbb{N}^+$ every $\mathcal{T}_n$ has height at most $\Delta$ and as $n \to \infty$ the minimal number of children of a nonleaf tends to infinity. We can view every tree as a (first-order) $\tau$-structure where $\tau$ is a signature with one binary relation symbol. For a fixed (arbitrary) finite and relational signature $\sigma \supseteq \tau$ we consider the set $\mathbf{W}_n$ of expansions of $\mathcal{T}_n$ to $\sigma$ and a probability distribution $\mathbb{P}_n$ on $\mathbf{W}_n$ which is determined by a (parametrized/lifted) Probabilistic Graphical Model (PGM) $\mathbb{G}$ which can use the information given by $\mathcal{T}_n$. The kind of PGM that we consider uses formulas of a many-valued logic that we call $PLA^*$ with truth values in the unit interval $[0, 1]$. We also use $PLA^*$ to express queries, or events, on $\mathbf{W}_n$. With this setup we prove that, under some assumptions on $\mathbf{T}$, $\mathbb{G}$, and a (possibly quite complex) formula $\varphi(x_1, \ldots, x_k)$ of $PLA^*$, as $n \to \infty$, if $a_1, \ldots, a_k$ are vertices of the tree $\mathcal{T}_n$ then the value of $\varphi(a_1, \ldots, a_k)$ will, with high probability, be almost the same as the value of $\psi(a_1, \ldots, a_k)$, where $\psi(x_1, \ldots, x_k)$ is a ``simple'' formula the value of which can always be computed quickly (without reference to $n$), and $\psi$ itself can be found by using only the information that defines $\mathbf{T}$, $\mathbb{G}$ and $\varphi$. A corollary of this, subject to the same conditions, is a probabilistic convergence law for $PLA^*$-formulas.

cs.LO

Random expansions of finite structures with bounded degree

We consider finite relational signatures $\tau \subseteq \sigma$, a sequence of finite base $\tau$-structures $(\mathcal{B}_n : n \in \mathbb{N})$ the cardinalities of which tend to infinity and such that, for some number $\Delta$, the degree of (the Gaifman graph of) every $\mathcal{B}_n$ is at most $\Delta$. We let $\mathbf{W}_n$ be the set of all expansions of $\mathcal{B}_n$ to $\sigma$ and we consider a probabilistic graphical model, a concept used in machine learning and artificial intelligence, to generate a probability distribution $\mathbb{P}_n$ on $\mathbf{W}_n$ for all $n$. We use a many-valued ``probability logic'' with truth values in the unit interval to express probabilities within probabilistic graphical models and to express queries on $\mathbf{W}_n$. This logic uses aggregation functions (e.g. the average) instead of quantifiers and it can express all queries (on finite structures) that can be expressed with first-order logic since the aggregation functions maximum and minimum can be used to express existential and universal quantifications, respectively. The main results concern asymptotic elimination of aggregation functions (the analogue of almost sure elimination of quantifiers for two-valued logics with quantifiers) and the asymptotic distribution of truth values of formulas, the analogue of logical convergence results for two-valued logics. The structure theory that is developed for sequences $(\mathcal{B}_n : n \in \mathbb{N})$ as above may be of independent interest.

math.LO

Asymptotic elimination of partially continuous aggregation functions in directed graphical models

In Statistical Relational Artificial Intelligence, a branch of AI and machine learning which combines the logical and statistical schools of AI, one uses the concept {\em para\-metrized probabilistic graphical model (PPGM)} to model (conditional) dependencies between random variables and to make probabilistic inferences about events on a space of "possible worlds". The set of possible worlds with underlying domain $D$ (a set of objects) can be represented by the set $\mathbf{W}_D$ of all first-order structures (for a suitable signature) with domain $D$. Using a formal logic we can describe events on $\mathbf{W}_D$. By combining a logic and a PPGM we can also define a probability distribution $\mathbb{P}_D$ on $\mathbf{W}_D$ and use it to compute the probability of an event. We consider a logic, denoted $PLA$, with truth values in the unit interval, which uses aggregation functions, such as arithmetic mean, geometric mean, maximum and minimum instead of quantifiers. However we face the problem of computational efficiency and this problem is an obstacle to the wider use of methods from Statistical Relational AI in practical applications. We address this problem by proving that the described probability will, under certain assumptions on the PPGM and the sentence $φ$, converge as the size of $D$ tends to infinity. The convergence result is obtained by showing that every formula $φ(x_1, \ldots, x_k)$ which contains only "admissible" aggregation functions (e.g. arithmetic and geometric mean, max and min) is asymptotically equivalent to a formula $ψ(x_1, \ldots, x_k)$ without aggregation functions.

cs.LO

A general approach to asymptotic elimination of aggregation functions and generalized quantifiers

We consider a logic with truth values in the unit interval and which uses aggregation functions instead of quantifiers, and we describe a general approach to asymptotic elimination of aggregation functions and, indirectly, of asymptotic elimination of Mostowski style generalized quantifiers, since such can be expressed by using aggregation functions. The notion of ``local continuity'' of an aggregation function, which we make precise in two (related) ways, plays a central role in this approach.

math.LO

On the relative asymptotic expressivity of inference frameworks

We consider logics with truth values in the unit interval $[0,1]$. Such logics are used to define queries and to define probability distributions. In this context the notion of almost sure equivalence of formulas is generalized to the notion of asymptotic equivalence. We prove two new results about the asymptotic equivalence of formulas where each result has a convergence law as a corollary. These results as well as several older results can be formulated as results about the relative asymptotic expressivity of inference frameworks. An inference framework $\mathbf{F}$ is a class of pairs $(\mathbb{P}, L)$, where $\mathbb{P} = (\mathbb{P}_n : n = 1, 2, 3, \ldots)$, $\mathbb{P}_n$ are probability distributions on the set $\mathbf{W}_n$ of all $\sigma$-structures with domain $\{1, \ldots, n\}$ (where $\sigma$ is a first-order signature) and $L$ is a logic with truth values in the unit interval $[0, 1]$. An inference framework $\mathbf{F}'$ is asymptotically at least as expressive as an inference framework $\mathbf{F}$ if for every $(\mathbb{P}, L) \in \mathbf{F}$ there is $(\mathbb{P}', L') \in \mathbf{F}'$ such that $\mathbb{P}$ is asymptotically total variation equivalent to $\mathbb{P}'$ and for every $\varphi(\bar{x}) \in L$ there is $\varphi'(\bar{x}) \in L'$ such that $\varphi'(\bar{x})$ is asymptotically equivalent to $\varphi(\bar{x})$ with respect to $\mathbb{P}$. This relation is a preorder. If, in addition, $\mathbf{F}$ is at least as expressive as $\mathbf{F}'$ then we say that $\mathbf{F}$ and $\mathbf{F}'$ are asymptotically equally expressive. Our third contribution is to systematize the new results of this paper and several previous results in order to get a preorder on a number of inference systems that are of relevance in the context of machine learning and artificial intelligence.

cs.LO

Conditional probability logic, lifted bayesian networks and almost sure quantifier elimination

We introduce a formal logical language, called conditional probability logic (CPL), which extends first-order logic and which can express probabilities, conditional probabilities and which can compare conditional probabilities. Intuitively speaking, although formal details are different, CPL can express the same kind of statements as some languages which have been considered in the artificial intelligence community. We also consider a way of making precise the notion of lifted Bayesian network, where this notion is a type of (lifted) probabilistic graphical model used in machine learning, data mining and artificial intelligence. A lifted Bayesian network (in the sense defined here) determines, in a natural way, a probability distribution on the set of all structures (in the sense of first-order logic) with a common finite domain $D$. Our main result is that for every "noncritical" CPL-formula $φ(\bar{x})$ there is a quantifier-free formula $φ^*(\bar{x})$ which is "almost surely" equivalent to $φ(\bar{x})$ as the cardinality of $D$ tends towards infinity. This is relevant for the problem of making probabilistic inferences on large domains $D$, because (a) the problem of evaluating, by "brute force", the probability of $φ(\bar{x})$ being true for some sequence $\bar{d}$ of elements from $D$ has, in general, (highly) exponential time complexity in the cardinality of $D$, and (b) the corresponding probability for the quantifier-free $φ^*(\bar{x})$ depends only on the lifted Bayesian network and not on $D$. The main result has two corollaries, one of which is a convergence law (and zero-one law) for noncritial CPL-formulas.

math.LO

On constraints and dividing in ternary homogeneous structures

Let M be ternary, homogeneous and simple. We prove that if M is finitely constrained, then it is supersimple with finite SU-rank and dependence is $k$-trivial for some $k < ω$ and for finite sets of real elements. Now suppose that, in addition, M is supersimple with SU-rank 1. If M is finitely constrained then algebraic closure in M is trivial. We also find connections between the nature of the constraints of M, the nature of the amalgamations allowed by the age of M, and the nature of definable equivalence relations. A key method of proof is to "extract" constraints (of M) from instances of dividing and from definable equivalence relations. Finally, we give new examples, including an uncountable family, of ternary homogeneous supersimple structures of SU-rank 1.

math.LO

Supersimple omega-categorical theories and pregeometries

We prove that if $T$ is an $ω$-categorical supersimple theory with nontrivial dependence (given by forking), then there is a nontrivial regular 1-type over a finite set of reals which is realized by real elements; hence forking induces a nontrivial pregeometry on the solution set of this type and the pregeometry is definable (using only finitely many parameters). The assumption about $ω$-categoricity is necessary. This result is used to prove the following: If $V$ is a finite relational vocabulary with maximal arity 3 and $T$ is a supersimple $V$-theory with elimination of quantifiers, then $T$ has trivial dependence and finite SU-rank. This immediately gives the following strengthening of a previous result of the author: if $\mathcal{M}$ is a ternary simple homogeneous structure with only finitely many constraints, then $Th(\mathcal{M})$ has trivial dependence and finite SU-rank.

math.LO

Binary simple homogeneous structures

We describe all binary simple homogeneous structures M in terms of 0-definable equivalence relations on M, which "coordinatize" M and control dividing, and extension properties that respect these equivalence relations.

math.LO

Binary primitive homogeneous simple structures

Suppose that M is countable, binary, primitive, homogeneous, and simple, and hence 1-based. We prove that the SU-rank of the complete theory of M is~1. It follows that M is a random structure. The conclusion that M is a random structure does not hold if the binarity condition is removed, as witnessed by the generic tetrahedron-free 3-hypergraph. However, to show that the generic tetrahedron-free 3-hypergraph is 1-based requires some work (it is known that it has the other properties) since this notion is defined in terms of imaginary elements. This is partly why we also characterize equivalence relations which are definable without parameters in the context of omega-categorical structures with degenerate algebraic closure. Another reason is that such characterizations may be useful in future research about simple (nonbinary) homogeneous structures.

math.LO

Homogeneous 1-based structures and interpretability in random structures

Let $V$ be a finite relational vocabulary in which no symbol has arity greater than 2. Let $M$ be countable $V$-structure which is homogeneous, simple and 1-based. The first main result says that if $M$ is, in addition, primitive, then it is strongly interpretable in a random structure. The second main result, which generalizes the first, implies (without the assumption on primitivity) that if $M$ is "coordinatized" by a set with SU-rank 1 and there is no definable (without parameters) nontrivial equivalence relation on $M$ with only finite classes, then $M$ is strongly interpretable in a random structure.

math.LO

Binary simple homogeneous structures are supersimple with finite rank

Suppose that M is an infinite structure with finite relational vocabulary such that every relation symbol has arity at most 2. If M is simple and homogeneous then its complete theory is supersimple with finite SU-rank which cannot exceed the number of complete 2-types over the empty set.

math.LO