SearcharxivSearch

arXiv subjects

Jeremy Levesley

Publications and source records attributed to Jeremy Levesley.

At least 19 recordsLinked to original sources

What is Hiding in Medicine's Dark Matter? Learning with Missing Data in Medical Practices

Electronic patient records (EPRs) produce a wealth of data but contain significant missing information. Understanding and handling this missing data is an important part of clinical data analysis and if left unaddressed could result in bias in analysis and distortion in critical conclusions. Missing data may be linked to health care professional practice patterns and imputation of missing data can increase the validity of clinical decisions. This study focuses on statistical approaches for understanding and interpreting the missing data and machine learning based clinical data imputation using a single centre's paediatric emergency data and the data from UK's largest clinical audit for traumatic injury database (TARN). In the study of 56,961 data points related to initial vital signs and observations taken on children presenting to an Emergency Department, we have shown that missing data are likely to be non-random and how these are linked to health care professional practice patterns. We have then examined 79 TARN fields with missing values for 5,791 trauma cases. Singular Value Decomposition (SVD) and k-Nearest Neighbour (kNN) based missing data imputation methods are used and imputation results against the original dataset are compared and statistically tested. We have concluded that the 1NN imputer is the best imputation which indicates a usual pattern of clinical decision making: find the most similar patients and take their attributes as imputation.

cs.LG

An Informational Space Based Semantic Analysis for Scientific Texts

One major problem in Natural Language Processing is the automatic analysis and representation of human language. Human language is ambiguous and deeper understanding of semantics and creating human-to-machine interaction have required an effort in creating the schemes for act of communication and building common-sense knowledge bases for the 'meaning' in texts. This paper introduces computational methods for semantic analysis and the quantifying the meaning of short scientific texts. Computational methods extracting semantic feature are used to analyse the relations between texts of messages and 'representations of situations' for a newly created large collection of scientific texts, Leicester Scientific Corpus. The representation of scientific-specific meaning is standardised by replacing the situation representations, rather than psychological properties, with the vectors of some attributes: a list of scientific subject categories that the text belongs to. First, this paper introduces 'Meaning Space' in which the informational representation of the meaning is extracted from the occurrence of the word in texts across the scientific categories, i.e., the meaning of a word is represented by a vector of Relative Information Gain about the subject categories. Then, the meaning space is statistically analysed for Leicester Scientific Dictionary-Core and we investigate 'Principal Components of the Meaning' to describe the adequate dimensions of the meaning. The research in this paper conducts the base for the geometric representation of the meaning of texts.

cs.CL

Convergence of sparse grid Gaussian convolution approximation for multi-dimensional periodic function

We consider the problem of approximating $[0,1]^{d}$-periodic functions by convolution with a scaled Gaussian kernel. We start by establishing convergence rates to functions from periodic Sobolev spaces and we show that the saturation rate is $O(h^{2}),$ where $h$ is the scale of the Gaussian kernel. Taken from a discrete point of view, this result can be interpreted as the accuracy that can be achieved on the uniform grid with spacing $h.$ In the discrete setting, the curse of dimensionality would place severe restrictions on the computation of the approximation. For instance, a spacing of $2^{-n}$ would provide an approximation converging at a rate of $O(2^{-2n})$ but would require $(2^{n}+1)^{d}$ grid points. To overcome this we introduce a sparse grid version of Gaussian convolution approximation, where substantially fewer grid points are required, and show that the sparse grid version delivers a saturation rate of $O(n^{d-1}2^{-2n}).$ This rate is in line with what one would expect in the sparse grid setting (where the full grid error only deteriorates by a factor of order $n^{d-1}$) however the analysis that leads to the result is novel in that it draws on results from the theory of special functions and key observations regarding the form of certain weighted geometric sums.

math.NA

Semantic Analysis for Automated Evaluation of the Potential Impact of Research Articles

Can the analysis of the semantics of words used in the text of a scientific paper predict its future impact measured by citations? This study details examples of automated text classification that achieved 80% success rate in distinguishing between highly-cited and little-cited articles. Automated intelligent systems allow the identification of promising works that could become influential in the scientific community. The problems of quantifying the meaning of texts and representation of human language have been clear since the inception of Natural Language Processing. This paper presents a novel method for vector representation of text meaning based on information theory and show how this informational semantics is used for text classification on the basis of the Leicester Scientific Corpus. We describe the experimental framework used to evaluate the impact of scientific articles through their informational semantics. Our interest is in citation classification to discover how important semantics of texts are in predicting the citation count. We propose the semantics of texts as an important factor for citation prediction. For each article, our system extracts the abstract of paper, represents the words of the abstract as vectors in Meaning Space, automatically analyses the distribution of scientific categories (Web of Science categories) within the text of abstract, and then classifies papers according to citation counts (highly-cited, little-cited). We show that an informational approach to representing the meaning of a text has offered a way to effectively predict the scientific impact of research papers.

cs.CL

Principal Components of the Meaning

In this paper we argue that (lexical) meaning in science can be represented in a 13 dimension Meaning Space. This space is constructed using principal component analysis (singular decomposition) on the matrix of word category relative information gains, where the categories are those used by the Web of Science, and the words are taken from a reduced word set from texts in the Web of Science. We show that this reduced word set plausibly represents all texts in the corpus, so that the principal component analysis has some objective meaning with respect to the corpus. We argue that 13 dimensions is adequate to describe the meaning of scientific texts, and hypothesise about the qualitative meaning of the principal components.

cs.CL

Convergence of Multilevel Stationary Gaussian Convolution

It is well-known that polynomial reproduction is not possible when approximating with Gaussian kernels. Quasi-interpolation schemes have been developed which use a finite number of Gaussians at different scales, which then reproduce polynomials of low degree \cite{beatson}, and thus achieve polynomial orders of convergence. At the same time, interpolation with kernels of fixed width suffers from an explosion in condition number, and information from all data points influences the approximation at any one data point (no localisation). In \cite{HL1} the authors show that, for periodic convolution with the Gaussian kernel, a multilevel scheme can give orders of approximation faster than any polynomial. In this paper we present a new multilevel quasi-interpolation algorithm, the discrete version of the algorithm in \cite{HL1}, which mimics the continuous algorithm well, to single precision accuracy, and gives excellent convergence rates for band limited periodic functions. In this paper we explain how the algorithm works, and why we achieve the numerical results we do. The estimates developed have two parts, one involving the convergence of a low degree polynomial truncation term and one involving the control of the remainder of the truncation as the algorithm proceeds.

math.NA

Personality Traits and Drug Consumption. A Story Told by Data

This is a preprint version of the first book from the series: "Stories told by data". In this book a story is told about the psychological traits associated with drug consumption. The book includes: - A review of published works on the psychological profiles of drug users. - Analysis of a new original database with information on 1885 respondents and usage of 18 drugs. (Database is available online.) - An introductory description of the data mining and machine learning methods used for the analysis of this dataset. - The demonstration that the personality traits (five factor model, impulsivity, and sensation seeking), together with simple demographic data, give the possibility of predicting the risk of consumption of individual drugs with sensitivity and specificity above 70% for most drugs. - The analysis of correlations of use of different substances and the description of the groups of drugs with correlated use (correlation pleiades). - Proof of significant differences of personality profiles for users of different drugs. This is explicitly proved for benzodiazepines, ecstasy, and heroin. - Tables of personality profiles for users and non-users of 18 substances. The book is aimed at advanced undergraduates or first-year PhD students, as well as researchers and practitioners. No previous knowledge of machine learning, advanced data mining concepts or modern psychology of personality is assumed. For more detailed introduction into statistical methods we recommend several undergraduate textbooks. Familiarity with basic statistics and some experience in the use of probabilities would be helpful as well as some basic technical understanding of psychology.

stat.AP

Automatic Short Answer Grading and Feedback Using Text Mining Methods

Automatic grading is not a new approach but the need to adapt the latest technology to automatic grading has become very important. As the technology has rapidly became more powerful on scoring exams and essays, especially from the 1990s onwards, partially or wholly automated grading systems using computational methods have evolved and have become a major area of research. In particular, the demand of scoring of natural language responses has created a need for tools that can be applied to automatically grade these responses. In this paper, we focus on the concept of automatic grading of short answer questions such as are typical in the UK GCSE system, and providing useful feedback on their answers to students. We present experimental results on a dataset provided from the introductory computer science class in the University of North Texas. We first apply standard data mining techniques to the corpus of student answers for the purpose of measuring similarity between the student answers and the model answer. This is based on the number of common words. We then evaluate the relation between these similarities and marks awarded by scorers. We then consider an approach that groups student answers into clusters. Each cluster would be awarded the same mark, and the same feedback given to each answer in a cluster. In this manner, we demonstrate that clusters indicate the groups of students who are awarded the same or the similar scores. Words in each cluster are compared to show that clusters are constructed based on how many and which words of the model answer have been used. The main novelty in this paper is that we design a model to predict marks based on the similarities between the student answers and the model answer.

cs.CL

Multilevel sparse grids collocation for linear partial differential equations, with tensor product smooth basis functions

Radial basis functions have become a popular tool for approximation and solution of partial differential equations (PDEs). The recently proposed multilevel sparse interpolation with kernels (MuSIK) algorithm proposed in \cite{Georgoulis} shows good convergence. In this paper we use a sparse kernel basis for the solution of PDEs by collocation. We will use the form of approximation proposed and developed by Kansa \cite{Kansa1986}. We will give numerical examples using a tensor product basis with the multiquadric (MQ) and Gaussian basis functions. This paper is novel in that we consider space-time PDEs in four dimensions using an easy-to-implement algorithm, with smooth approximations. The accuracy observed numerically is as good, with respect to the number of data points used, as other methods in the literature; see \cite{Langer1,Wang1}.

math.NA

Pseudo-Outcrop Visualization of Borehole Images and Core Scans

A pseudo-outcrop visualization is demonstrated for borehole and full-diameter rock core images to augment the ubiquitous unwrapped cylinder view and thereby to assist non-specialist interpreters. The pseudo-outcrop visualization is equivalent to a nonlinear projection of the image from borehole to earth frame of reference that creates a solid volume sliced longitudinally to reveal two or more faces in which the orientations of geological features indicate what is observed in the subsurface. A proxy for grain size is used to modulate the external dimensions of the plot to mimic profiles seen in real outcrops. The volume is created from a mixture of geological boundary elements and texture, the latter being the residue after the sum of boundary elements is subtracted from the original data. In the case of measurements from wireline microresistivity tools, whose circumferential coverage is substantially less than 100%, the missing circumferential data is first inpainted using multiscale directional transforms, which decompose the image into its elemental building structures, before reconstructing the full image. The pseudo-outcrop view enables direct observation of the angular relationships between features and aids visual comparison between borehole and core images, especially for the interested non-specialist.

physics.geo-ph

Convergence of Multilevel Stationary Gaussian Quasi-Interpolation

In this paper we present a new multilevel quasi-interpolation algorithm for smooth periodic functions using scaled Gaussians as basis functions. Recent research in this area has focussed upon implementations using basis function with finite smoothness. In this paper we deliver a first error estimates for the multilevel algorithm using analytic basis functions. The estimate has two parts, one involving the convergence of a low degree polynomial truncation term and one involving the control of the remainder of the truncation as the algorithm proceeds. Thus, numerically one observes a convergent scheme. Numerical results suggest that the scheme converges much faster than the theory shows.

math.NA

Quasi-interpolation on a sparse grid with Gaussian

Motivated by the recent multilevel sparse kernel-based interpolation (MuSIK) algorithm proposed in [Georgoulis, Levesley and Subhan, SIAM J. Sci. Comput., 35(2), pp. A815-A831, 2013], we introduce the new quasi-multilevel sparse interpolation with kernels (Q-MuSIK) via the combination technique. The Q-MuSIK scheme achieves better convergence and run time in comparison with classical quasi-interpolation; namely, the Q-MuSIK algorithm is generally superior to the MuSIK methods in terms of run time in particular in high-dimensional interpolation problems, since there is no need to solve large algebraic systems. We subsequently propose a fast, low complexity, high-dimensional quadrature formula based on Q-MuSIK interpolation of the integrand. We present the results of numerical experimentation for both interpolation and quadrature in high dimension.

math.NA

Approximation of exponential-type functions on a uniform grid by shifts of a basis function

In this paper, we study the problem of interpolating a continuous function at $(n+1)$ equally-spaced points in the interval $[0,1]$, using shifts of a kernel on the $(1/n)$-spaced infinite grid. The archetypal example here is approximation using shifts of a Gaussian kernel. We present new results concerning interpolation of functions of exponential type, in particular, polynomials on the integer grid as a step en route to solve the general interpolation problem. For the Gaussian kernel we introduce a new class of polynomials, closely related to the probabilistic Hermite polynomials and show that evaluations of the polynomials at the integer points provide the coefficients of the interpolants. Taking cue from the classical Newton polynomial interpolation, we derive a closed formula for the Gaussian interpolant of a continuous function on a uniform grid in the unit interval.

math.NA

Fast multilevel sparse Gaussian kernels for high-dimensional approximation and integration

A fast multilevel algorithm based on directionally scaled tensor-product Gaussian kernels on structured sparse grids is proposed for interpolation of high-dimensional functions and for the numerical integration of high-dimensional integrals. The algorithm is based on the recent Multilevel Sparse Kernel-based Interpolation (MLSKI) method (Georgoulis, Levesley \& Subhan, \emph{SIAM J. Sci. Comput.}, 35(2), pp.~A815--A831, 2013), with particular focus on the fast implementation of Gaussian-based MLSKI for interpolation and integration problems of high-dimen-sional functions $f:[0,1]^d\to\mathbb{R}$, with $5\le d\le 10$. The MLSKI interpolation procedure is shown to be interpolatory and a fast implementation is proposed. More specifically, exploiting the tensor-product nature of anisotropic Gaussian kernels, one-dimensional cardinal basis functions on a sequence of hierarchical equidistant nodes are precomputed to machine precision, rendering the interpolation problem into a fully parallelisable ensemble of linear combinations of function evaluations. A numerical integration algorithm is also proposed, based on interpolating the (high-dimensional) integrand. A series of numerical experiments highlights the applicability of the proposed algorithm for interpolation and integration for up to 10-dimensional problems.

math.NA

Magnetic Flux Leakage Method: Large-Scale Approximation

We consider the application of the magnetic flux leakage (MFL) method to the detection of defects in ferromagnetic (steel) tubulars. The problem setup corresponds to the cases where the distance from the casing and the point where the magnetic field is measured is small compared to the curvature radius of the undamaged casing and the scale of inhomogeneity of the magnetic field in the defect-free case. Mathematically this corresponds to the planar ferromagnetic layer in a uniform magnetic field oriented along this layer. Defects in the layer surface result in a strong deformation of the magnetic field, which provides opportunities for the reconstruction of the surface profile from measurements of the magnetic field. We deal with large-scale defects whose depth is small compared to their longitudinal sizes---these being typical of corrosive damage. Within the framework of large-scale approximation, analytical relations between the casing thickness profile and the measured magnetic field can be derived.

physics.data-an

Noise-Produced Patterns in Images Constructed from Magnetic Flux Leakage Data

Magnetic flux leakage measurements help identify the position, size and shape of corrosion-related defects in steel casings used to protect boreholes drilled into oil and gas reservoirs. Images constructed from magnetic flux leakage data contain patterns related to noise inherent in the method. We investigate the patterns and their scaling properties for the case of delta-correlated input noise, and consider the implications for the method's ability to resolve defects. The analytical evaluation of the noise-produced patterns is made possible by model reduction facilitated by large-scale approximation. With appropriate modification, the approach can be employed to analyze noise-produced patterns in other situations where the data of interest are not measured directly, but are related to the measured data by a complex linear transform involving integrations with respect to spatial coordinates.

physics.data-an

Lévy driven models and derivative pricing

We develop a general method for derivative pricing. This approach has its roots in Shannon's Information Theory. The notion of $λ$-analyticity of Lévy models is introduced on the basis of which new representations of the pricing integral are obtained. It is shown that popular in applications Lévy models are $λ$-analytic. We apply these results to derive a general algorithm for pricing of European call options.

stat.AP

On the density of polyharmonic splines

This article treats the question of fundamentality of the translates of a polyharmonic spline kernel (also known as a surface spline) in the space of continuous functions on a compact set $Ω\subset \RR^d$ when the translates are restricted to $Ω$. Fundamentality is not hard to demonstrate when a low degree polynomial may be added or when translates are permitted to lie outside of $Ω$; the challenge of this problem stems from the presence of the boundary, for which all successful approximation schemes require an added polynomial. When $Ω$ is the unit ball, we demonstrate that translates of polyharmonic splines are fundamental by considering two related problems: the fundamentality in the space of functions vanishing at the boundary and fundamentality of the restricted kernel in the space of continuous function on the sphere. This gives rise to a new approximation scheme composed of two parts: one which approximates purely on $\partial Ω$, and a second part involving a shift invariant approximant of a function vanishing outside of a neighborhood $Ω$.

math.CA