SearcharxivSearch

arXiv subjects

Tadas Temčinas

Publications and source records attributed to Tadas Temčinas.

4 recordsLinked to original sources

Goodness-of-fit via Count Statistics in Dense Random Simplicial Complexes

A key object of study in stochastic topology is a random simplicial complex. In this work we study a multi-parameter random simplicial complex model, where the probability of including a $k$-simplex, given the lower dimensional structure, is fixed. This leads to a conditionally independent probabilistic structure. This model includes the Erdős-Rényi random graph, the random clique complex as well as the Linial-Meshulam complex as special cases. The model is studied from both probabilistic and statistical points of view. We prove multivariate central limit theorems with bounds and known limiting covariance structure for the subcomplex counts and the number of critical simplices under a lexicographical acyclic partial matching. We use the CLTs to develop a goodness-of-fit test for this random model and evaluate its empirical performance. In order for the test to be applicable in practice, we also prove that the MLE estimators are asymptotically unbiased, consistent, uncorrelated and normally distributed.

math.ST

Multivariate Central Limit Theorems for Random Clique Complexes

Motivated by open problems in applied and computational algebraic topology, we establish multivariate normal approximation theorems for three random vectors which arise organically in the study of random clique complexes. These are: (1) the vector of critical simplex counts attained by a lexicographical Morse matching, (2) the vector of simplex counts in the link of a fixed simplex, and (3) the vector of total simplex counts. The first of these random vectors forms a cornerstone of modern homology algorithms, while the second one provides a natural generalisation for the notion of vertex degree, and the third one may be viewed from the perspective of U-statistics. To obtain distributional approximations for these random vectors, we extend the notion of dissociated sums to a multivariate setting and prove a new central limit theorem for such sums using Stein's method.

math.PR

Efficient Intent Detection with Dual Sentence Encoders

Building conversational systems in new domains and with added functionality requires resource-efficient models that work under low-data regimes (i.e., in few-shot setups). Motivated by these requirements, we introduce intent detection methods backed by pretrained dual sentence encoders such as USE and ConveRT. We demonstrate the usefulness and wide applicability of the proposed intent detectors, showing that: 1) they outperform intent detectors based on fine-tuning the full BERT-Large model or using BERT as a fixed black-box encoder on three diverse intent detection data sets; 2) the gains are especially pronounced in few-shot setups (i.e., with only 10 or 30 annotated examples per intent); 3) our intent detectors can be trained in a matter of minutes on a single CPU; and 4) they are stable across different hyperparameter settings. In hope of facilitating and democratizing research focused on intention detection, we release our code, as well as a new challenging single-domain intent detection dataset comprising 13,083 annotated examples over 77 intents.

cs.CL

Local Homology of Word Embeddings

Topological data analysis (TDA) has been widely used to make progress on a number of problems. However, it seems that TDA application in natural language processing (NLP) is at its infancy. In this paper we try to bridge the gap by arguing why TDA tools are a natural choice when it comes to analysing word embedding data. We describe a parallelisable unsupervised learning algorithm based on local homology of datapoints and show some experimental results on word embedding data. We see that local homology of datapoints in word embedding data contains some information that can potentially be used to solve the word sense disambiguation problem.

math.AT