SearcharxivSearch

arXiv subjects

Emma Slade

Publications and source records attributed to Emma Slade.

At least 19 recordsLinked to original sources

Out-of-distribution evaluations of channel agnostic masked autoencoders in fluorescence microscopy

Developing computer vision for high-content screening is challenging due to various sources of distribution-shift caused by changes in experimental conditions, perturbagens, and fluorescent markers. The impact of different sources of distribution-shift are confounded in typical evaluations of models based on transfer learning, which limits interpretations of how changes to model design and training affect generalisation. We propose an evaluation scheme that isolates sources of distribution-shift using the JUMP-CP dataset, allowing researchers to evaluate generalisation with respect to specific sources of distribution-shift. We then present a channel-agnostic masked autoencoder $\mathbf{Campfire}$ which, via a shared decoder for all channels, scales effectively to datasets containing many different fluorescent markers, and show that it generalises to out-of-distribution experimental batches, perturbagens, and fluorescent markers, and also demonstrates successful transfer learning from one cell type to another.

cs.LG

Dataset Distillation as Pushforward Optimal Quantization

Dataset distillation aims to find a synthetic training set such that training on the synthetic data achieves similar performance to training on real data, with orders of magnitude less computational requirements. Existing methods can be broadly categorized as either bi-level optimization problems that have neural network training heuristics as the lower level problem, or disentangled methods that bypass the bi-level optimization by matching distributions of data. The latter method has the major advantages of speed and scalability in terms of size of both training and distilled datasets. We demonstrate that when equipped with an encoder-decoder structure, the empirically successful disentangled methods can be reformulated as an optimal quantization problem, where a finite set of points is found to approximate the underlying probability measure by minimizing the expected projection distance. In particular, we link existing disentangled dataset distillation methods to the classical optimal quantization and Wasserstein barycenter problems, demonstrating consistency of distilled datasets for diffusion-based generative priors. We propose Dataset Distillation by Optimal Quantization, based on clustering in a latent space. Compared to the previous SOTA method D\textsuperscript{4}M, we achieve better performance and inter-model generalization on the ImageNet-1K dataset with trivial additional computation, and SOTA performance in higher image-per-class settings. Using the distilled noise initializations in a stronger diffusion transformer model, we obtain SOTA distillation performance on ImageNet-1K and its subsets, outperforming diffusion guidance methods.

cs.LG

Self-supervised learning of multi-omics embeddings in the low-label, high-data regime

Contrastive, self-supervised learning (SSL) is used to train a model that predicts cancer type from miRNA, mRNA or RPPA expression data. This model, a pretrained FT-Transformer, is shown to outperform XGBoost and CatBoost, standard benchmarks for tabular data, when labelled samples are scarce but the number of unlabelled samples is high. This is despite the fact that the datasets we use have $\mathcal{O}(10^{1})$ classes and $\mathcal{O}(10^{2})-\mathcal{O}(10^{4})$ features. After demonstrating the efficacy of our chosen method of self-supervised pretraining, we investigate SSL for multi-modal models. A late-fusion model is proposed, where each omics is passed through its own sub-network, the outputs of which are averaged and passed to the pretraining or downstream objective function. Multi-modal pretraining is shown to improve predictions from a single omics, and we argue that this is useful for datasets with many unlabelled multi-modal samples, but few labelled unimodal samples. Additionally, we show that pretraining each omics-specific module individually is highly effective. This enables the application of the proposed model in a variety of contexts where a large amount of unlabelled data is available from each omics, but only a few labelled samples.

cs.LG

Mining of Single-Class by Active Learning for Semantic Segmentation

Several Active Learning (AL) policies require retraining a target model several times in order to identify the most informative samples and rarely offer the option to focus on the acquisition of samples from underrepresented classes. Here the Mining of Single-Class by Active Learning (MiSiCAL) paradigm is introduced where an AL policy is constructed through deep reinforcement learning and exploits quantity-accuracy correlations to build datasets on which high-performance models can be trained with regards to specific classes. MiSiCAL is especially helpful in the case of very large batch sizes since it does not require repeated model training sessions as is common in other AL methods. This is thanks to its ability to exploit fixed representations of the candidate data points. We find that MiSiCAL is able to outperform a random policy on 150 out of 171 COCO10k classes, while the strongest baseline only outperforms random on 101 classes.

cs.LG

Deep reinforced active learning for multi-class image classification

High accuracy medical image classification can be limited by the costs of acquiring more data as well as the time and expertise needed to label existing images. In this paper, we apply active learning to medical image classification, a method which aims to maximise model performance on a minimal subset from a larger pool of data. We present a new active learning framework, based on deep reinforcement learning, to learn an active learning query strategy to label images based on predictions from a convolutional neural network. Our framework modifies the deep-Q network formulation, allowing us to pick data based additionally on geometric arguments in the latent space of the classifier, allowing for high accuracy multi-class classification in a batch-based active learning setting, enabling the agent to label datapoints that are both diverse and about which it is most uncertain. We apply our framework to two medical imaging datasets and compare with standard query strategies as well as the most recent reinforcement learning based active learning approach for image classification.

cs.CV

GNisi: A graph network for reconstructing Ising models from multivariate binarized data

Ising models are a simple generative approach to describing interacting binary variables. They have proven useful in a number of biological settings because they enable one to represent observed many-body correlations as the separable consequence of many direct, pairwise statistical interactions. The inference of Ising models from data can be computationally very challenging and often one must be satisfied with numerical approximations or limited precision. In this paper we present a novel method for the determination of Ising parameters from data, called GNisi, which uses a Graph Neural network trained on known Ising models in order to construct the parameters for unseen data. We show that GNisi is more accurate than the existing state of the art software, and we illustrate our method by applying GNisi to gene expression data.

cs.LG

Data efficiency in graph networks through equivariance

We introduce a novel architecture for graph networks which is equivariant to any transformation in the coordinate embeddings that preserves the distance between neighbouring nodes. In particular, it is equivariant to the Euclidean and conformal orthogonal groups in $n$-dimensions. Thanks to its equivariance properties, the proposed model is extremely more data efficient with respect to classical graph architectures and also intrinsically equipped with a better inductive bias. We show that, learning on a minimal amount of data, the architecture we propose can perfectly generalise to unseen data in a synthetic problem, while much more training data are required from a standard model to reach comparable performance.

cs.LG

Cuts for two-body decays at colliders

Fixed-order perturbative calculations of fiducial cross sections for two-body decay processes at colliders show disturbing sensitivity to unphysically low momentum scales and, in the case of $H\to \gamma \gamma$ in gluon fusion, poor convergence. Such problems have their origins in an interplay between the behaviour of standard experimental cuts at small transverse momenta ($p_t$) and logarithmic perturbative contributions. We illustrate how this interplay leads to a factorially divergent structure in the perturbative series that sets in already from the first orders. We propose simple modifications of fiducial cuts to eliminate their key incriminating characteristic, a linear dependence of the acceptance on the Higgs or $Z$-boson $p_t$, replacing it with quadratic dependence. This brings major improvements in the behaviour of the perturbative expansion. More elaborate cuts can achieve an acceptance that is independent of the Higgs $p_t$ at low $p_t$, with a variety of consequent advantages.

hep-ph

Symmetry-driven graph neural networks

Exploiting symmetries and invariance in data is a powerful, yet not fully exploited, way to achieve better generalisation with more efficiency. In this paper, we introduce two graph network architectures that are equivariant to several types of transformations affecting the node coordinates. First, we build equivariance to any transformation in the coordinate embeddings that preserves the distance between neighbouring nodes, allowing for equivariance to the Euclidean group. Then, we introduce angle attributes to build equivariance to any angle preserving transformation - thus, to the conformal group. Thanks to their equivariance properties, the proposed models can be vastly more data efficient with respect to classical graph architectures, intrinsically equipped with a better inductive bias and better at generalising. We demonstrate these capabilities on a synthetic dataset composed of $n$-dimensional geometric objects. Additionally, we provide examples of their limitations when (the right) symmetries are not present in the data.

cs.LG

Combined SMEFT interpretation of Higgs, diboson, and top quark data from the LHC

We present a global interpretation of Higgs, diboson, and top quark production and decay measurements from the LHC in the framework of the Standard Model Effective Field Theory (SMEFT) at dimension six. We constrain simultaneously 36 independent directions in its parameter space, and compare the outcome of the global analysis with that from individual and two-parameter fits. Our results are obtained by means of state-of-the-art theoretical calculations for the SM and the EFT cross-sections, and account for both linear and quadratic corrections in the $1/\Lambda^2$ expansion. We demonstrate how the inclusion of NLO QCD and $\mathcal{O}\left( \Lambda^{-4}\right)$ effects is instrumental to accurately map the posterior distributions associated to the fitted Wilson coefficients. We assess the interplay and complementarity between the top quark, Higgs, and diboson measurements, deploy a variety of statistical estimators to quantify the impact of each dataset in the parameter space, and carry out fits in BSM-inspired scenarios such as the top-philic model. Our results represent a stepping stone in the ongoing program of model-independent searches at the LHC from precision measurements, and pave the way towards yet more global SMEFT interpretations extended to other high-$p_T$ processes as well as to low-energy observables.

hep-ph

Beyond permutation equivariance in graph networks

In this draft paper, we introduce a novel architecture for graph networks which is equivariant to the Euclidean group in $n$-dimensions. The model is designed to work with graph networks in their general form and can be shown to include particular variants as special cases. Thanks to its equivariance properties, we expect the proposed model to be more data efficient with respect to classical graph architectures and also intrinsically equipped with a better inductive bias. We defer investigating this matter to future work.

cs.LG

Computing Tools for the SMEFT

The increasing interest in the phenomenology of the Standard Model Effective Field Theory (SMEFT), has led to the development of a wide spectrum of public codes which implement automatically different aspects of the SMEFT for phenomenological applications. In order to discuss the present and future of such efforts, the "SMEFT-Tools 2019" Workshop was held at the IPPP Durham on the 12th-14th June 2019. Here we collect and summarize the contents of this workshop.

hep-ph

Towards global fits in EFT's and New Physics implications

I discuss recent progress on fits to dimension-six operators in the Standard Model Effective Theory (SMEFT). I focus on the top quark sector of the SMEFT, as well as the theoretical advances made in computing SMEFT effects through to next-to-leading order in QCD and the use of these calculations in global fits. I also discuss fits performed to the Higgs and electroweak sectors of the SMEFT and the possibility for performing global fits to multiple sectors simultaneously.

hep-ph

Constraining the SMEFT with Bayesian reweighting

We illustrate how Bayesian reweighting can be used to incorporate the constraints provided by new measurements into a global Monte Carlo analysis of the Standard Model Effective Field Theory (SMEFT). This method, extensively applied to study the impact of new data on the parton distribution functions of the proton, is here validated by means of our recent SMEFiT analysis of the top quark sector. We show how, under well-defined conditions and for the SMEFT operators directly sensitive to the new data, the reweighting procedure is equivalent to a corresponding new fit. We quantify the amount of information added to the SMEFT parameter space by means of the Shannon entropy and of the Kolmogorov-Smirnov statistic. We investigate the dependence of our results upon the choice of either the NNPDF or the Giele-Keller expressions of the weights.

hep-ph

A Monte Carlo analysis of the SMEFT in the top quark sector

We present a framework for carrying out global analyses of the Standard Model Effective Field Theory: SMEFiT. This approach is based on the Monte Carlo replica method, widely used in the case of NNPDF fits of the proton structure, for deriving a faithful estimate of the experimental and theoretical uncertainties. As a proof of concept of the SMEFiT methodology, we present a study of the constraints on the SMEFT provided by top quark production measurements from the LHC. We derive bounds for the 34 degrees of freedom relevant for the interpretation of the LHC top quark data and compare these bounds with previously reported constraints.

hep-ph

A Monte Carlo global analysis of the Standard Model Effective Field Theory: the top quark sector

We present a novel framework for carrying out global analyses of the Standard Model Effective Field Theory (SMEFT) at dimension-six: SMEFiT. This approach is based on the Monte Carlo replica method for deriving a faithful estimate of the experimental and theoretical uncertainties and enables one to construct the probability distribution in the space of the SMEFT degrees of freedom. As a proof of concept of the SMEFiT methodology, we present a first study of the constraints on the SMEFT provided by top quark production measurements from the LHC. Our analysis includes more than 30 independent measurements from 10 different processes at 8 and 13 TeV such as inclusive top-quark pair and single-top production and the associated production of top quarks with weak vector bosons and the Higgs boson. State-of-the-art theoretical calculations are adopted both for the Standard Model and for the SMEFT contributions, where in the latter case NLO QCD corrections are included for the majority of processes. We derive bounds for the 34 degrees of freedom relevant for the interpretation of the LHC top quark data and compare these bounds with previously reported constraints. Our study illustrates the significant potential of LHC precision measurements to constrain physics beyond the Standard Model in a model-independent way, and paves the way towards a global analysis of the SMEFT.

hep-ph

Precision determination of the strong coupling constant within a global PDF analysis

We present a determination of the strong coupling constant $\alpha_s(m_Z)$ based on the NNPDF3.1 determination of parton distributions, which for the first time includes constraints from jet production, top-quark pair differential distributions, and the $Z$ $p_T$ distributions using exact NNLO theory. Our result is based on a novel extension of the NNPDF methodology - the correlated replica method - which allows for a simultaneous determination of $\alpha_s$ and the PDFs with all correlations between them fully taken into account. We study in detail all relevant sources of experimental, methodological and theoretical uncertainty. At NNLO we find $\alpha_s(m_Z) = 0.1185 \pm 0.0005^\text{(exp)}\pm 0.0001^\text{(meth)}$, showing that methodological uncertainties are negligible. We conservatively estimate the theoretical uncertainty due to missing higher order QCD corrections (N$^3$LO and beyond) from half the shift between the NLO and NNLO $\alpha_s$ values, finding $\Delta\alpha^{\rm th}_s =0.0011$.

hep-ph

Direct photon production and PDF fits reloaded

Direct photon production in hadronic collisions provides a handle on the gluon PDF by means of the QCD Compton scattering process. In this work we revisit the impact of direct photon production on a global PDF analysis, motivated by the recent availability of the next-to-next-to-leading (NNLO) calculation for this process. We demonstrate that the inclusion of NNLO QCD and leading-logarithmic electroweak corrections leads to a good quantitative agreement with the ATLAS measurements at 8 TeV and 13 TeV, except for the most forward rapidity region in the former case. By including the ATLAS 8 TeV direct photon production data in the NNPDF3.1 NNLO global analysis, we assess its impact on the medium-x gluon. We also study the constraining power of the direct photon production measurements on PDF fits based on different datasets, in particular on the NNPDF3.1 no-LHC and collider-only fits. We also present updated NNLO theoretical predictions for direct photon production at 13 TeV that include the constraints from the 8 TeV measurements.

hep-ph