Searcharxiv⌕ Search

arXiv subjects

Dominik Janzing

Publications and source records attributed to Dominik Janzing.

At least 91 records · Page 5Linked to original sources

Identifiability of Causal Graphs using Functional Models

This work addresses the following question: Under what assumptions on the data generating process can one infer the causal graph from the joint distribution? The approach taken by conditional independence-based causal discovery methods is based on two assumptions: the Markov condition and faithfulness. It has been shown that under these assumptions the causal graph can be identified up to Markov equivalence (some arrows remain undirected) using methods like the PC algorithm. In this work we propose an alternative by defining Identifiable Functional Model Classes (IFMOCs). As our main theorem we prove that if the data generating process belongs to an IFMOC, one can identify the complete causal graph. To the best of our knowledge this is the first identifiability result of this kind that is not limited to linear functional relationships. We discuss how the IFMOC assumption and the Markov and faithfulness assumptions relate to each other and explain why we believe that the IFMOC assumption can be tested more easily on given data. We further provide a practical algorithm that recovers the causal graph from finitely many data; experiments on simulated data support the theoretical findings.

cs.LG↗

Kernel-based Conditional Independence Test and Application in Causal Discovery

Conditional independence testing is an important problem, especially in Bayesian network learning and causal discovery. Due to the curse of dimensionality, testing for conditional independence of continuous variables is particularly challenging. We propose a Kernel-based Conditional Independence test (KCI-test), by constructing an appropriate test statistic and deriving its asymptotic distribution under the null hypothesis of conditional independence. The proposed method is computationally efficient and easy to implement. Experimental results show that it outperforms other methods, especially when the conditioning set is large or the sample size is not very large, in which case other methods encounter difficulties.

cs.LG↗

Testing whether linear equations are causal: A free probability theory approach

We propose a method that infers whether linear relations between two high-dimensional variables X and Y are due to a causal influence from X to Y or from Y to X. The earlier proposed so-called Trace Method is extended to the regime where the dimension of the observed variables exceeds the sample size. Based on previous work, we postulate conditions that characterize a causal relation between X and Y. Moreover, we describe a statistical test and argue that both causal directions are typically rejected if there is a common cause. A full theoretical analysis is presented for the deterministic case but our approach seems to be valid for the noisy case, too, for which we additionally present an approach based on a sparsity constraint. The discussed method yields promising results for both simulated and real world data.

cs.LG↗

Robust Learning via Cause-Effect Models

We consider the problem of function estimation in the case where the data distribution may shift between training and test time, and additional information about it may be available at test time. This relates to popular scenarios such as covariate shift, concept drift, transfer learning and semi-supervised learning. This working paper discusses how these tasks could be tackled depending on the kind of changes of the distributions. It argues that knowledge of an underlying causal direction can facilitate several of these tasks.

stat.ML↗

Thermodynamic limits of dynamic cooling

We study dynamic cooling, where an externally driven two-level system is cooled via reservoir, a quantum system with initial canonical equilibrium state. We obtain explicitly the minimal possible temperature $T_{\rm min}>0$ reachable for the two-level system. The minimization goes over all unitary dynamic processes operating on the system and reservoir, and over the reservoir energy spectrum. The minimal work needed to reach $T_{\rm min}$ grows as $1/T_{\rm min}$. This work cost can be significantly reduced, though, if one is satisfied by temperatures slightly above $T_{\rm min}$. Our results on $T_{\rm min}>0$ prove unattainability of the absolute zero temperature without ambiguities that surround its derivation from the entropic version of the third law. The unattainability can be recovered, albeit via a different mechanism, for cooling by a reservoir with an initially microcanonic state. We also study cooling via a reservoir consisting of $N\gg 1$ identical spins. Here we show that $T_{\rm min}\propto\frac{1}{N}$ and find the maximal cooling compatible with the minimal work determined by the free energy.

cond-mat.stat-mech↗

Is there a physically universal cellular automaton or Hamiltonian?

It is known that both quantum and classical cellular automata (CA) exist that are computationally universal in the sense that they can simulate, after appropriate initialization, any quantum or classical computation, respectively. Here we introduce a different notion of universality: a CA is called physically universal if every transformation on any finite region can be (approximately) implemented by the autonomous time evolution of the system after the complement of the region has been initialized in an appropriate way. We pose the question of whether physically universal CAs exist. Such CAs would provide a model of the world where the boundary between a physical system and its controller can be consistently shifted, in analogy to the Heisenberg cut for the quantum measurement problem. We propose to study the thermodynamic cost of computation and control within such a model because implementing a cyclic process on a microsystem may require a non-cyclic process for its controller, whereas implementing a cyclic process on system and controller may require the implementation of a non-cyclic process on a "meta"-controller, and so on. Physically universal CAs avoid this infinite hierarchy of controllers and the cost of implementing cycles on a subsystem can be described by mixing properties of the CA dynamics. We define a physical prior on the CA configurations by applying the dynamics to an initial state where half of the CA is in the maximum entropy state and half of it is in the all-zero state (thus reflecting the fact that life requires non-equilibrium states like the boundary between a hold and a cold reservoir). As opposed to Solomonoff's prior, our prior does not only account for the Kolmogorov complexity but also for the cost of isolating the system during the state preparation if the preparation process is not robust.

quant-ph↗

Causal Markov condition for submodular information measures

The causal Markov condition (CMC) is a postulate that links observations to causality. It describes the conditional independences among the observations that are entailed by a causal hypothesis in terms of a directed acyclic graph. In the conventional setting, the observations are random variables and the independence is a statistical one, i.e., the information content of observations is measured in terms of Shannon entropy. We formulate a generalized CMC for any kind of observations on which independence is defined via an arbitrary submodular information measure. Recently, this has been discussed for observations in terms of binary strings where information is understood in the sense of Kolmogorov complexity. Our approach enables us to find computable alternatives to Kolmogorov complexity, e.g., the length of a text after applying existing data compression schemes. We show that our CMC is justified if one restricts the attention to a class of causal mechanisms that is adapted to the respective information measure. Our justification is similar to deriving the statistical CMC from functional models of causality, where every variable is a deterministic function of its observed causes and an unobserved noise term. Our experiments on real data demonstrate the performance of compression based causal inference.

cs.IT↗

Causal Inference on Discrete Data using Additive Noise Models

Inferring the causal structure of a set of random variables from a finite sample of the joint distribution is an important problem in science. Recently, methods using additive noise models have been suggested to approach the case of continuous variables. In many situations, however, the variables of interest are discrete or even have only finitely many states. In this work we extend the notion of additive noise models to these cases. We prove that whenever the joint distribution $\prob^{(X,Y)}$ admits such a model in one direction, e.g. $Y=f(X)+N, N \independent X$, it does not admit the reversed model $X=g(Y)+\tilde N, \tilde N \independent Y$ as long as the model is chosen in a generic way. Based on these deliberations we propose an efficient new algorithm that is able to distinguish between cause and effect for a finite sample of discrete variables. In an extensive experimental study we show that this algorithm works both on synthetic and real data sets.

stat.ML↗

Distinguishing Cause and Effect via Second Order Exponential Models

We propose a method to infer causal structures containing both discrete and continuous variables. The idea is to select causal hypotheses for which the conditional density of every variable, given its causes, becomes smooth. We define a family of smooth densities and conditional densities by second order exponential models, i.e., by maximizing conditional entropy subject to first and second statistical moments. If some of the variables take only values in proper subsets of R^n, these conditionals can induce different families of joint distributions even for Markov-equivalent graphs. We consider the case of one binary and one real-valued variable where the method can distinguish between cause and effect. Using this example, we describe that sometimes a causal hypothesis must be rejected because P(effect|cause) and P(cause) share algorithmic information (which is untypical if they are chosen independently). This way, our method is in the same spirit as faithfulness-based causal inference because it also rejects non-generic mutual adjustments among DAG-parameters.

stat.ML↗

Justifying additive-noise-model based causal discovery via algorithmic information theory

A recent method for causal discovery is in many cases able to infer whether X causes Y or Y causes X for just two observed variables X and Y. It is based on the observation that there exist (non-Gaussian) joint distributions P(X,Y) for which Y may be written as a function of X up to an additive noise term that is independent of X and no such model exists from Y to X. Whenever this is the case, one prefers the causal model X--> Y. Here we justify this method by showing that the causal hypothesis Y--> X is unlikely because it requires a specific tuning between P(Y) and P(X|Y) to generate a distribution that admits an additive noise model from X to Y. To quantify the amount of tuning required we derive lower bounds on the algorithmic information shared by P(Y) and P(X|Y). This way, our justification is consistent with recent approaches for using algorithmic information theory for causal reasoning. We extend this principle to the case where P(X,Y) almost admits an additive noise model. Our results suggest that the above conclusion is more reliable if the complexity of P(Y) is high.

cs.IT↗

Telling cause from effect based on high-dimensional observations

We describe a method for inferring linear causal relations among multi-dimensional variables. The idea is to use an asymmetry between the distributions of cause and effect that occurs if both the covariance matrix of the cause and the structure matrix mapping cause to the effect are independently chosen. The method works for both stochastic and deterministic causal relations, provided that the dimensionality is sufficiently high (in some experiments, 5 was enough). It is applicable to Gaussian as well as non-Gaussian data.

stat.ML↗

On the entropy production of time series with unidirectional linearity

There are non-Gaussian time series that admit a causal linear autoregressive moving average (ARMA) model when regressing the future on the past, but not when regressing the past on the future. The reason is that, in the latter case, the regression residuals are only uncorrelated but not statistically independent of the future. In previous work, we have experimentally verified that many empirical time series indeed show such a time inversion asymmetry. For various physical systems, it is known that time-inversion asymmetries are linked to the thermodynamic entropy production in non-equilibrium states. Here we show that such a link also exists for the above unidirectional linearity. We study the dynamical evolution of a physical toy system with linear coupling to an infinite environment and show that the linearity of the dynamics is inherited to the forward-time conditional probabilities, but not to the backward-time conditionals. The reason for this asymmetry between past and future is that the environment permanently provides particles that are in a product state before they interact with the system, but show statistical dependencies afterwards. From a coarse-grained perspective, the interaction thus generates entropy. We quantitatively relate the strength of the non-linearity of the backward conditionals to the minimal amount of entropy generation.

cond-mat.stat-mech↗

Thermodynamic efficiency of information and heat flow

A basic task of information processing is information transfer (flow). Here we study a pair of Brownian particles each coupled to a thermal bath at temperature $T_1$ and $T_2$, respectively. The information flow in such a system is defined via the time-shifted mutual information. The information flow nullifies at equilibrium, and its efficiency is defined as the ratio of flow over the total entropy production in the system. For a stationary state the information flows from higher to lower temperatures, and its the efficiency is bound from above by $\frac{{\rm max}[T_1,T_2]}{|T_1-T_2|}$. This upper bound is imposed by the second law and it quantifies the thermodynamic cost for information flow in the present class of systems. It can be reached in the adiabatic situation, where the particles have widely different characteristic times. The efficiency of heat flow|defined as the heat flow over the total amount of dissipated heat|is limited from above by the same factor. There is a complementarity between heat- and information-flow: the setup which is most efficient for the former is the least efficient for the latter and {\it vice versa}. The above bound for the efficiency can be [transiently] overcome in certain non-stationary situations, but the efficiency is still limited from above. We study yet another measure of information-processing [transfer entropy] proposed in literature. Though this measure does not require any thermodynamic cost, the information flow and transfer entropy are shown to be intimately related for stationary states.

cond-mat.stat-mech↗

On causally asymmetric versions of Occam's Razor and their relation to thermodynamics

In real-life statistical data, it seems that conditional probabilities for the effect given their causes tend to be less complex and smoother than conditionals for causes, given their effects. We have recently proposed and tested methods for causal inference in machine learning using a formalization of this principle. Here we try to provide some theoretical justification for causal inference methods based upon such a ``causally asymmetric'' interpretation of Occam's Razor. To this end, we discuss toy models of cause-effect relations from classical and quantum physics as well as computer science in the context of various aspects of complexity. We argue that this asymmetry of the statistical dependences between cause and effect has a thermodynamic origin. The essential link is the tendency of the environment to provide independent background noise realized by physical systems that are initially uncorrelated with the system under consideration rather than being finally uncorrelated. This link extends ideas from the literature relating Reichenbach's principle of the common cause to the second law.

cond-mat.stat-mech↗

Causal inference using the algorithmic Markov condition

Inferring the causal structure that links n observables is usually based upon detecting statistical dependences and choosing simple graphs that make the joint measure Markovian. Here we argue why causal inference is also possible when only single observations are present. We develop a theory how to generate causal graphs explaining similarities between single objects. To this end, we replace the notion of conditional stochastic independence in the causal Markov condition with the vanishing of conditional algorithmic mutual information and describe the corresponding causal inference rules. We explain why a consistent reformulation of causal inference in terms of algorithmic complexity implies a new inference principle that takes into account also the complexity of conditional probability densities, making it possible to select among Markov equivalent causal graphs. This insight provides a theoretical foundation of a heuristic principle proposed in earlier work. We also discuss how to replace Kolmogorov complexity with decidable complexity criteria. This can be seen as an algorithmic analog of replacing the empirically undecidable question of statistical independence with practical independence tests that are based on implicit or explicit assumptions on the underlying distribution.

math.ST↗

A PromiseBQP-complete String Rewriting Problem

We are given three strings s, t, and t' of length L over some fixed finite alphabet and an integer m that is polylogarithmic in L. We have a symmetric relation on substrings of constant length that specifies which substrings are allowed to be replaced with each other. Let Delta(n) denote the difference between the numbers of possibilities to obtain t from s and t' from s after n replacements. The problem is to determine the sign of Delta(m). As promises we have a gap condition and a growth condition. The former states that |Delta(m)| >= epsilon c^m where epsilon is inverse polylogarithmic in L and c>0 is a constant. The latter is given by Delta(n) <= c^n for all n. We show that this problem is PromiseBQP-complete, i.e., it represents the class of problems which can be solved efficiently on a quantum computer.

quant-ph↗

A single-shot measurement of the energy of product states in a translation invariant spin chain can replace any quantum computation

In measurement-based quantum computation, quantum algorithms are implemented via sequences of measurements. We describe a translationally invariant finite-range interaction on a one-dimensional qudit chain and prove that a single-shot measurement of the energy of an appropriate computational basis state with respect to this Hamiltonian provides the output of any quantum circuit. The required measurement accuracy scales inverse polynomially with the size of the simulated quantum circuit. This shows that the implementation of energy measurements on generic qudit chains is as hard as the realization of quantum computation. Here a ''measurement'' is any procedure that samples from the spectral measure induced by the observable and the state under consideration. As opposed to measurement-based quantum computation, the post-measurement state is irrelevant.

quant-ph↗

How much is a quantum controller controlled by the controlled system?

We consider unitary transformations on a bipartite system A x B. To what extent entails the ability to transmit information from A to B the ability to transfer information in the converse direction? We prove a dimension-dependent lower bound on the classical channel capacity C(A<--B) in terms of the capacity C(A-->B) for the case that the bipartite unitary operation consists of controlled local unitaries on B conditioned on basis states on A. This can be interpreted as a statement on the strength of the inevitable backaction of a quantum system on its controller. If the local operations are given by the regular representation of a finite group G we have C(A-->B)=log |G| and C(A<--B)=log N where N is the sum over the degrees of all inequivalent representations. Hence the information deficit C(A-->B)-C(A<--B) between the forward and the backward capacity depends on the "non-abelianness" of the control group. For regular representations, the ratio between backward and forward capacities cannot be smaller than 1/2. The symmetric group S_n reaches this bound asymptotically. However, for the general case (without group structure) all bounds must depend on the dimensions since it is known that the ratio can tend to zero.

quant-ph↗