SearcharxivSearch

arXiv subjects

Francesco Caltagirone

Publications and source records attributed to Francesco Caltagirone.

At least 19 recordsLinked to original sources

Conditioned Text Generation with Transfer for Closed-Domain Dialogue Systems

Scarcity of training data for task-oriented dialogue systems is a well known problem that is usually tackled with costly and time-consuming manual data annotation. An alternative solution is to rely on automatic text generation which, although less accurate than human supervision, has the advantage of being cheap and fast. Our contribution is twofold. First we show how to optimally train and control the generation of intent-specific sentences using a conditional variational autoencoder. Then we introduce a new protocol called query transfer that allows to leverage a large unlabelled dataset, possibly containing irrelevant queries, to extract relevant information. Comparison with two different baselines shows that this method, in the appropriate regime, consistently improves the diversity of the generated queries without compromising their quality. We also demonstrate the effectiveness of our generation method as a data augmentation technique for language modelling tasks.

cs.CL

Conditioned Query Generation for Task-Oriented Dialogue Systems

Scarcity of training data for task-oriented dialogue systems is a well known problem that is usually tackled with costly and time-consuming manual data annotation. An alternative solution is to rely on automatic text generation which, although less accurate than human supervision, has the advantage of being cheap and fast. In this paper we propose a novel controlled data generation method that could be used as a training augmentation framework for closed-domain dialogue. Our contribution is twofold. First we show how to optimally train and control the generation of intent-specific sentences using a conditional variational autoencoder. Then we introduce a novel protocol called query transfer that allows to leverage a broad, unlabelled dataset to extract relevant information. Comparison with two different baselines shows that our method, in the appropriate regime, consistently improves the diversity of the generated queries without compromising their quality.

cs.CL

Spoken Language Understanding on the Edge

We consider the problem of performing Spoken Language Understanding (SLU) on small devices typical of IoT applications. Our contributions are twofold. First, we outline the design of an embedded, private-by-design SLU system and show that it has performance on par with cloud-based commercial solutions. Second, we release the datasets used in our experiments in the interest of reproducibility and in the hope that they can prove useful to the SLU community.

cs.CL

Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces

This paper presents the machine learning architecture of the Snips Voice Platform, a software solution to perform Spoken Language Understanding on microprocessors typical of IoT devices. The embedded inference is fast and accurate while enforcing privacy by design, as no personal user data is ever collected. Focusing on Automatic Speech Recognition and Natural Language Understanding, we detail our approach to training high-performance Machine Learning models that are small enough to run in real-time on small devices. Additionally, we describe a data generation procedure that provides sufficient, high-quality training data without compromising user privacy.

cs.CL

A Deterministic and Generalized Framework for Unsupervised Learning with Restricted Boltzmann Machines

Restricted Boltzmann machines (RBMs) are energy-based neural-networks which are commonly used as the building blocks for deep architectures neural architectures. In this work, we derive a deterministic framework for the training, evaluation, and use of RBMs based upon the Thouless-Anderson-Palmer (TAP) mean-field approximation of widely-connected systems with weak interactions coming from spin-glass theory. While the TAP approach has been extensively studied for fully-visible binary spin systems, our construction is generalized to latent-variable models, as well as to arbitrarily distributed real-valued spin systems with bounded support. In our numerical experiments, we demonstrate the effective deterministic training of our proposed models and are able to show interesting features of unsupervised learning which could not be directly observed with sampling. Additionally, we demonstrate how to utilize our TAP-based framework for leveraging trained RBMs as joint priors in denoising problems.

cs.LG

Recovering asymmetric communities in the stochastic block model

We consider the sparse stochastic block model in the case where the degrees are uninformative. The case where the two communities have approximately the same size has been extensively studied and we concentrate here on the community detection problem in the case of unbalanced communities. In this setting, spectral algorithms based on the non-backtracking matrix are known to solve the community detection problem (i.e. do strictly better than a random guess) when the signal is sufficiently large namely above the so-called Kesten Stigum threshold. In this regime and when the average degree tends to infinity, we show that if the community of a vanishing fraction of the vertices is revealed, then a local algorithm (belief propagation) is optimal down to Kesten Stigum threshold and we quantify explicitly its performance. Below the Kesten Stigum threshold, we show that, in the large degree limit, there is a second threshold called the spinodal curve below which, the community detection problem is not solvable. The spinodal curve is equal to the Kesten Stigum threshold when the fraction of vertices in the smallest community is above $p^*=\frac{1}{2}-\frac{1}{2\sqrt{3}}$, so that the Kesten Stigum threshold is the threshold for solvability of the community detection in this case. However when the smallest community is smaller than $p^*$, the spinodal curve only provides a lower bound on the threshold for solvability. In the regime below the Kesten Stigum bound and above the spinodal curve, we also characterize the performance of best local algorithms as a function of the fraction of revealed vertices. Our proof relies on a careful analysis of the associated reconstruction problem on trees which might be of independent interest. In particular, we show that the spinodal curve corresponds to the reconstruction threshold on the tree.

math.PR

Blind Sensor Calibration using Approximate Message Passing

The ubiquity of approximately sparse data has led a variety of com- munities to great interest in compressed sensing algorithms. Although these are very successful and well understood for linear measurements with additive noise, applying them on real data can be problematic if imperfect sensing devices introduce deviations from this ideal signal ac- quisition process, caused by sensor decalibration or failure. We propose a message passing algorithm called calibration approximate message passing (Cal-AMP) that can treat a variety of such sensor-induced imperfections. In addition to deriving the general form of the algorithm, we numerically investigate two particular settings. In the first, a fraction of the sensors is faulty, giving readings unrelated to the signal. In the second, sensors are decalibrated and each one introduces a different multiplicative gain to the measures. Cal-AMP shares the scalability of approximate message passing, allowing to treat big sized instances of these problems, and ex- perimentally exhibits a phase transition between domains of success and failure.

cs.IT

Inferring Sparsity: Compressed Sensing using Generalized Restricted Boltzmann Machines

In this work, we consider compressed sensing reconstruction from $M$ measurements of $K$-sparse structured signals which do not possess a writable correlation model. Assuming that a generative statistical model, such as a Boltzmann machine, can be trained in an unsupervised manner on example signals, we demonstrate how this signal model can be used within a Bayesian framework of signal reconstruction. By deriving a message-passing inference for general distribution restricted Boltzmann machines, we are able to integrate these inferred signal models into approximate message passing for compressed sensing reconstruction. Finally, we show for the MNIST dataset that this approach can be very effective, even for $M < K$.

cs.IT

Random Projections through multiple optical scattering: Approximating kernels at the speed of light

Random projections have proven extremely useful in many signal processing and machine learning applications. However, they often require either to store a very large random matrix, or to use a different, structured matrix to reduce the computational and memory costs. Here, we overcome this difficulty by proposing an analog, optical device, that performs the random projections literally at the speed of light without having to store any matrix in memory. This is achieved using the physical properties of multiple coherent scattering of coherent light in random media. We use this device on a simple task of classification with a kernel machine, and we show that, on the MNIST database, the experimental results closely match the theoretical performance of the corresponding kernel. This framework can help make kernel methods practical for applications that have large training sets and/or require real-time prediction. We discuss possible extensions of the method in terms of a class of kernels, speed, memory consumption and different problems.

cs.ET

Spectral Detection on Sparse Hypergraphs

We consider the problem of the assignment of nodes into communities from a set of hyperedges, where every hyperedge is a noisy observation of the community assignment of the adjacent nodes. We focus in particular on the sparse regime where the number of edges is of the same order as the number of vertices. We propose a spectral method based on a generalization of the non-backtracking Hashimoto matrix into hypergraphs. We analyze its performance on a planted generative model and compare it with other spectral methods and with Bayesian belief propagation (which was conjectured to be asymptotically optimal for this model). We conclude that the proposed spectral method detects communities whenever belief propagation does, while having the important advantages to be simpler, entirely nonparametric, and to be able to learn the rule according to which the hyperedges were generated without prior information.

cs.SI

Replica Theory and Spin Glasses

These are notes from the lectures of Giorgio Parisi given at the autumn school "Statistical Physics, Optimization, Inference, and Message-Passing Algorithm", that took place in Les Houches, France from Monday September 30th, 2013, till Friday October 11th, 2013. The school was organized by Florent Krzakala from UPMC and ENS Paris, Federico Ricci-Tersenghi from "La Sapienza" Roma, Lenka Zdeborová from CEA Saclay and CNRS, and Riccardo Zecchina from Politecnico Torino. The first lecture contains an introduction to the replica method, along with a concrete application to the computation of the eigenvalue distribution of random matrices in the GOE. In the second lecture, the solution of the SK model is derived, along with the phenomenon of replica symmetry breaking (RSB). In the third part, the physical meaning of the RSB is explained. The ultrametricity of the space of pure states emerges as a consequence of the hierarchical RSB scheme. Moreover, it is shown how some low temperature properties of physical observables can be derived by invoking the stochastic stability principle. Lecture four contains some rigorous results on the SK model: the existence of the thermodynamic limit, and the proof of the exactness of the hierarchical RSB solution.

cond-mat.stat-mech

Properties of spatial coupling in compressed sensing

In this paper we address a series of open questions about the construction of spatially coupled measurement matrices in compressed sensing. For hardware implementations one is forced to depart from the limiting regime of parameters in which the proofs of the so-called threshold saturation work. We investigate quantitatively the behavior under finite coupling range, the dependence on the shape of the coupling interaction, and optimization of the so-called seed to minimize distance from optimality. Our analysis explains some of the properties observed empirically in previous works and provides new insight on spatially coupled compressed sensing.

cs.IT

On Convergence of Approximate Message Passing

Approximate message passing is an iterative algorithm for compressed sensing and related applications. A solid theory about the performance and convergence of the algorithm exists for measurement matrices having iid entries of zero mean. However, it was observed by several authors that for more general matrices the algorithm often encounters convergence problems. In this paper we identify the reason of the non-convergence for measurement matrices with iid entries and non-zero mean in the context of Bayes optimal inference. Finally we demonstrate numerically that when the iterative update is changed from parallel to sequential the convergence is restored.

cs.IT

Dynamics and termination cost of spatially coupled mean-field models

This work is motivated by recent progress in information theory and signal processing where the so-called `spatially coupled' design of systems leads to considerably better performance. We address relevant open questions about spatially coupled systems through the study of a simple Ising model. In particular, we consider a chain of Curie-Weiss models that are coupled by interactions up to a certain range. Indeed, it is well known that the pure (uncoupled) Curie-Weiss model undergoes a first order phase transition driven by the magnetic field, and furthermore, in the spinodal region such systems are unable to reach equilibrium in sub-exponential time if initialized in the metastable state. By contrast, the spatially coupled system is, instead, able to reach the equilibrium even when initialized to the metastable state. The equilibrium phase propagates along the chain in the form of a travelling wave. Here we study the speed of the wave-front and the so-called `termination cost'--- \textit{i.e.}, the conditions necessary for the propagation to occur. We reach several interesting conclusions about optimization of the speed and the cost.

cond-mat.stat-mech

Blind Calibration in Compressed Sensing using Message Passing Algorithms

Compressed sensing (CS) is a concept that allows to acquire compressible signals with a small number of measurements. As such it is very attractive for hardware implementations. Therefore, correct calibration of the hardware is a central is- sue. In this paper we study the so-called blind calibration, i.e. when the training signals that are available to perform the calibration are sparse but unknown. We extend the approximate message passing (AMP) algorithm used in CS to the case of blind calibration. In the calibration-AMP, both the gains on the sensors and the elements of the signals are treated as unknowns. Our algorithm is also applica- ble to settings in which the sensors distort the measurements in other ways than multiplication by a gain, unlike previously suggested blind calibration algorithms based on convex relaxations. We study numerically the phase diagram of the blind calibration problem, and show that even in cases where convex relaxation is pos- sible, our algorithm requires a smaller number of measurements and/or signals in order to perform well.

cs.IT

Critical Off-Equilibrium Dynamics in Glassy Systems

We consider off-equilibrium dynamics at the critical temperature in a class of glassy system. The off-equilibrium correlation and response functions obey a precise scaling form in the aging regime. The structure of the {\it equilibrium} replicated Gibbs free energy fixes the corresponding {\it off-equilibrium} scaling functions implicitly through two functional equations. The details of the model enter these equations only through the ratio $w_2/w_1$ of the cubic coefficients (proper vertexes) of the replicated Gibbs free energy. Therefore the off-equilibrium dynamical exponents are controlled by the very same parameter exponent $λ=w_2/w_1$ that determines equilibrium dynamics. We find approximate solutions to the equations and validate the theory by means of analytical computations and numerical simulations.

cond-mat.dis-nn

Dynamical critical exponents for the mean-field Potts glass

In this paper we study the critical behaviour of the fully-connected p-colours Potts model at the dynamical transition. In the framework of Mode Coupling Theory (MCT), the time autocorrelation function displays a two step relaxation, with two exponents governing the approach to the plateau and the exit from it. Exploiting a relation between statics and equilibrium dynamics which has been recently introduced, we are able to compute the critical slowing down exponents at the dynamical transition with arbitrary precision and for any value of the number of colours p. When available, we compare our exact results with numerical simulations. In addition, we present a detailed study of the dynamical transition in the large p limit, showing that the system is not equivalent to a random energy model.

cond-mat.dis-nn

Critical slowing down exponents in quenched disordered spin models for structural glasses: Random Orthogonal and related models

An important prediction of Mode-Coupling-Theory (MCT) is the relationship between the power- law decay exponents in the β regime. In the original structural glass context this relationship follows from the MCT equations that are obtained making rather uncontrolled approximations and λ has to be treated like a tunable parameter. It is known that a certain class of mean-field spin-glass models is exactly described by MCT equations. In this context, the physical meaning of the so called parameter exponent λ has recently been unveiled, giving a method to compute it exactly in a static framework. In this paper we exploit this new technique to compute the critical slowing down exponents in a class of mean-field Ising spin-glass models including, as special cases, the Sherrington-Kirkpatrick model, the p-spin model and the Random Orthogonal model.

cond-mat.dis-nn