SearcharxivSearch

arXiv subjects

Oliver Johnson

Publications and source records attributed to Oliver Johnson.

At least 19 recordsLinked to original sources

The Double Emulator

Computer models (simulators) are vital tools for investigating physical processes. Despite their utility, the prohibitive run-time of simulators hinders their direct application for uncertainty quantification. Gaussian process emulators (GPEs) have been used extensively to circumvent the cost of the simulator and are known to perform well on simulators with smooth, stationary output. In reality, many simulators violate these assumptions. Motivated by a finite element simulator which models early stage corrosion of uranium in water vapor, we propose an adaption of the GPE, called the double emulator, specifically for simulators which 'ground' in a considerable volume of their input space. Grounding is the process by which a simulator attains its minimum and can result in violation of the stationarity and smoothness assumptions used in the conventional GPE. We perform numerical experiments comparing the performance of the GPE and double emulator on both the corrosion simulator and synthetic examples.

stat.CO

Finite de Finetti bounds in relative entropy

We review old and recent finite de Finetti theorems in total variation distance and in relative entropy, and we highlight their connections with bounds on the difference between sampling with and without replacement. We also establish two new finite de Finetti theorems for exchangeable random vectors taking values in arbitrary spaces. These bounds are tight, and they are independent of the size and the dimension of the underlying space.

math.PR

Relative entropy bounds for sampling with and without replacement

Sharp, nonasymptotic bounds are obtained for the relative entropy between the distributions of sampling with and without replacement from an urn with balls of $c\geq 2$ colors. Our bounds are asymptotically tight in certain regimes and, unlike previous results, they depend on the number of balls of each colour in the urn. The connection of these results with finite de Finetti-style theorems is explored, and it is observed that a sampling bound due to Stam (1978) combined with the convexity of relative entropy yield a new finite de Finetti bound in relative entropy, which achieves the optimal asymptotic convergence rate.

math.PR

Small error algorithms for tropical group testing

We consider a version of the classical group testing problem motivated by PCR testing for COVID-19. In the so-called tropical group testing model, the outcome of a test is the lowest cycle threshold (Ct) level of the individuals pooled within it, rather than a simple binary indicator variable. We introduce the tropical counterparts of three classical non-adaptive algorithms (COMP, DD and SCOMP), and analyse their behaviour through both simulations and bounds on error probabilities. By comparing the results of the tropical and classical algorithms, we gain insight into the extra information provided by learning the outcomes (Ct levels) of the tests. We show that in a limiting regime the tropical COMP algorithm requires as many tests as its classical counterpart, but that for sufficiently dense problems tropical DD can recover more information with fewer tests, and can be viewed as essentially optimal in certain regimes.

cs.IT

Saved You A Click: Automatically Answering Clickbait Titles

Often clickbait articles have a title that is phrased as a question or vague teaser that entices the user to click on the link and read the article to find the explanation. We developed a system that will automatically find the answer or explanation of the clickbait hook from the website text so that the user does not need to read through the text themselves. We fine-tune an extractive question and answering model (RoBERTa) and an abstractive one (T5), using data scraped from the 'StopClickbait' Facebook pages and Reddit's 'SavedYouAClick' subforum. We find that both extractive and abstractive models improve significantly after finetuning. We find that the extractive model performs slightly better according to ROUGE scores, while the abstractive one has a slight edge in terms of BERTscores.

cs.CL

Extensions on `A Convex Scheme for the Secrecy Capacity of a MIMO Wiretap Channel with a Single Antenna Eavesdropper'

One key metric for physical layer security is the secrecy capacity. This is the maximum rate that a system can transmit with perfect secrecy. For a Multiple Input Multiple Output (MIMO) system (a newer technology for 5G, 6G and beyond) the secrecy capacity is not fully understood. For a Gaussian MIMO channel, the secrecy capacity is a non-convex optimisation problem for which a general solution is not available. Previous work by the authors showed that the secrecy capacity of a MIMO system with a single eavesdrop antenna is concave to a cut off point. In this work, which extends the previous paper, results are given for the region beyond this cut off point. It is shown that, for certain parameters, the presented scheme is concave to a point, and convex beyond it, and can therefore be solved efficiently using existing convex optimisation software.

cs.IT

A negative binomial approximation in group testing

We consider the problem of group testing (pooled testing), first introduced by Dorfman. For non-adaptive testing strategies, we refer to a non-defective item as `intruding' if it only appears in positive tests. Such items cause mis-classification errors in the well-known COMP algorithm, and can make other algorithms produce an error. It is therefore of interest to understand the distribution of the number of intruding items. We show that, under Bernoulli matrix designs, this distribution is well approximated in a variety of senses by a negative binomial distribution, allowing us to understand the performance of the two-stage conservative group testing algorithm of Aldridge.

math.PR

Bounds on Eavesdropper Performance for a MIMO-NOMA Downlink Scheme

Non-Orthogonal Multiple Access (NOMA) is a multiplexing technique for future wireless, which when combined with Multiple-Input Multiple-Output (MIMO) unlocks higher capacities for systems where users have varying channel strength. NOMA utilises the channel differences to increase the throughput, while MIMO exploits the additional degrees of freedom (DoF) to enhance this. This work analyses the secrecy capacity, demonstrating the robustness of a combined MIMO-NOMA scheme at physical layer, when in the presence of a passive eavesdropper. We present bounds on the eavesdropper performance and show heuristically that, as the number of users and antennas increases, the eavesdropper's SINR becomes small, regardless of how `lucky' they may be with their channel.

cs.IT

The COVID-19 pandemic as experienced by the individual

The ongoing COVID-19 pandemic has progressed with varying degrees of intensity in individual countries, suggesting it is important to analyse factors that vary between them. We study measures of `population-weighted density', which capture density as perceived by a randomly chosen individual. These measures of population density can significantly explain variation in the initial rate of spread of COVID-19 between countries within Europe. However, such measures do not explain differences on a global scale, particularly when considering countries in East Asia, or looking later into the epidemics. Therefore, to control for country-level differences in response to COVID-19 we consider the cross-cultural measure of individualism proposed by Hofstede. This score can significantly explain variation in the size of epidemics across Europe, North America, and East Asia. Using both our measure of population-weighted density and the Hofstede score we can significantly explain half the variation in the current size of epidemics across Europe and North America. By controlling for country-level responses to the virus and population density, our analysis of the global incidence of COVID-19 can help focus attention on epidemic control measures that are effective for individual countries.

physics.soc-ph

Information-theoretic convergence of extreme values to the Gumbel distribution

We show how convergence to the Gumbel distribution in an extreme value setting can be understood in an information-theoretic sense. We introduce a new type of score function which behaves well under the maximum operation, and which implies simple expressions for entropy and relative entropy. We show that, assuming certain properties of the von Mises representation, convergence to the Gumbel can be proved in the strong sense of relative entropy.

math.ST

Improved bounds for noisy group testing with constant tests per item

The group testing problem is concerned with identifying a small set of infected individuals in a large population. At our disposal is a testing procedure that allows us to test several individuals together. In an idealized setting, a test is positive if and only if at least one infected individual is included and negative otherwise. Significant progress was made in recent years towards understanding the information-theoretic and algorithmic properties in this noiseless setting. In this paper, we consider a noisy variant of group testing where test results are flipped with certain probability, including the realistic scenario where sensitivity and specificity can take arbitrary values. Using a test design where each individual is assigned to a fixed number of tests, we derive explicit algorithmic bounds for two commonly considered inference algorithms and thereby naturally extend the results of Scarlett \& Cevher (2016) and Scarlett \& Johnson (2020). We provide improved performance guarantees for the efficient algorithms in these noisy group testing models -- indeed, for a large set of parameter choices the bounds provided in the paper are the strongest currently proved.

cs.IT

Maximal correlation and the rate of Fisher information convergence in the Central Limit Theorem

We consider the behaviour of the Fisher information of scaled sums of independent and identically distributed random variables in the Central Limit Theorem regime. We show how this behaviour can be related to the second-largest non-trivial eigenvalue associated with the Hirschfeld--Gebelein--R\'{e}nyi maximal correlation. We prove that assuming this eigenvalue satisfies a strict inequality, an $O(1/n)$ rate of convergence and a strengthened form of monotonicity hold.

cs.IT

Group Testing: An Information Theory Perspective

The group testing problem concerns discovering a small number of defective items within a large population by performing tests on pools of items. A test is positive if the pool contains at least one defective, and negative if it contains no defectives. This is a sparse inference problem with a combinatorial flavour, with applications in medical testing, biology, telecommunications, information technology, data science, and more. In this monograph, we survey recent developments in the group testing problem from an information-theoretic perspective. We cover several related developments: efficient algorithms with practical storage and computation requirements, achievability bounds for optimal decoding methods, and algorithm-independent converse bounds. We assess the theoretical guarantees not only in terms of scaling laws, but also in terms of the constant factors, leading to the notion of the {\em rate} of group testing, indicating the amount of information learned per test. For the noiseless setting, we present a series of results leading to optimal rates, which in turn imply optimality and suboptimality results of various algorithms depending on the sparsity regime. We also survey analogous developments in noisy settings. In addition, we survey results concerning a number of variations on the standard group testing problem, including approximate recovery criteria, adaptive algorithms with a limited number of stages, sublinear-time algorithms, and settings with additional prior information, among others.

cs.IT

A proof of the Shepp-Olkin entropy monotonicity conjecture

Consider tossing a collection of coins, each fair or biased towards heads, and take the distribution of the total number of heads that result. It is natural to conjecture that this distribution should be 'more random' when each coin is fairer. Indeed, Shepp and Olkin conjectured that the Shannon entropy of this distribution is monotonically increasing in this case. We resolve this conjecture, by proving that this intuition is correct. Our proof uses a construction which was previously developed by the authors to prove a related conjecture of Shepp and Olkin concerning concavity of entropy. We discuss whether this result can be generalized to $q$-Rényi and $q$-Tsallis entropies, for a range of values of $q$.

math.PR

Reliability of Broadcast Communications Under Sparse Random Linear Network Coding

Ultra-reliable Point-to-Multipoint (PtM) communications are expected to become pivotal in networks offering future dependable services for smart cities. In this regard, sparse Random Linear Network Coding (RLNC) techniques have been widely employed to provide an efficient way to improve the reliability of broadcast and multicast data streams. This paper addresses the pressing concern of providing a tight approximation to the probability of a user recovering a data stream protected by this kind of coding technique. In particular, by exploiting the Stein--Chen method, we provide a novel and general performance framework applicable to any combination of system and service parameters, such as finite field sizes, lengths of the data stream and level of sparsity. The deviation of the proposed approximation from Monte Carlo simulations is negligible, improving significantly on the state of the art performance bounds.

cs.IT

Noisy Non-Adaptive Group Testing: A (Near-)Definite Defectives Approach

The group testing problem consists of determining a small set of defective items from a larger set of items based on a number of possibly-noisy tests, and is relevant in applications such as medical testing, communication protocols, pattern matching, and more. We study the noisy version of this problem, where the outcome of each standard noiseless group test is subject to independent noise, corresponding to passing the noiseless result through a binary channel. We introduce a class of algorithms that we refer to as Near-Definite Defectives (NDD), and study bounds on the required number of tests for asymptotically vanishing error probability under Bernoulli random test designs. In addition, we study algorithm-independent converse results, giving lower bounds on the required number of tests under Bernoulli test designs. Under reverse Z-channel noise, the achievable rates and converse results match in a broad range of sparsity regimes, and under Z-channel noise, the two match in a narrower range of dense/low-noise regimes. We observe that although these two channels have the same Shannon capacity when viewed as a communication channel, they can behave quite differently when it comes to group testing. Finally, we extend our analysis of these noise models to a general binary noise model (including symmetric noise), and show improvements over known existing bounds in broad scaling regimes.

cs.IT