SearcharxivSearch

arXiv subjects

Melissa Humphries

Publications and source records attributed to Melissa Humphries.

9 recordsLinked to original sources

Best Preprocessing Techniques for Sentiment Analysis

Sentiment analysis in Twitter datasets is important because it enables monitoring public opinion on products and analysis of political and social movements. One critical step is preprocessing: the automated processing of text for machine learning algorithms. Preprocessing plays a critical role in reducing noise and improving efficiency. However, little research has systematically examined the order in which preprocessing techniques are implemented. We find that, when accounting for order, spelling correction is the least impactful preprocessing technique, whereas tokenisation is the most impactful. Stemming and stop-word removal are interchangeable, and it is better to remove stop words without removing negation. The best order for applying the preprocessing techniques was tokenisation, text cleaning, stemming, and then stopword removal. Our results provide a systematic approach for practitioners to deploy preprocessing to improve model output without the costly preprocessing exploratory phase.

cs.CL

Repeatability is not recovery: Quantifying algorithmic stability and topic recovery in Latent Dirichlet Allocation

Topic models are often judged by the consistency of their outputs across repeated runs, implicitly assuming that repeatable topic output is a successful recovery of the underlying topics. We show that this assumption is false: repeatability is not recovery. We introduce a stability framework that jointly measures consistency among repeated runs and accuracy relative to known ground truth. Because real-world corpora lack known topic structures, we generate synthetic corpora using the Latent Dirichlet Allocation (LDA) generative process, enabling direct evaluation of topic recovery. Across 50 repeated LDA runs on each corpus, we find that LDA reliably identifies the correct number of topics and frequently converges to highly consistent topic solutions. However, these repeatable solutions frequently fail to recover the true generating topics. Thus, internal stability should not be interpreted as evidence of correctness. Our results illustrate that stability and recovery are distinct properties of topic models and should be evaluated separately. Consequently, topic-model outputs should be validated using multiple complementary criteria before supporting substantive conclusions, particularly in high-stakes applications.

cs.CL

A novel application of Shapley values for large multidimensional time-series data: Applying explainable AI to a DNA profile classification neural network

The application of Shapley values to high-dimensional, time-series-like data is computationally challenging - and sometimes impossible. For $N$ inputs the problem is $2^N$ hard. In image processing, clusters of pixels, referred to as superpixels, are used to streamline computations. This research presents an efficient solution for time-seres-like data that adapts the idea of superpixels for Shapley value computation. Motivated by a forensic DNA classification example, the method is applied to multivariate time-series-like data whose features have been classified by a convolutional neural network (CNN). In DNA processing, it is important to identify alleles from the background noise created by DNA extraction and processing. A single DNA profile has $31,200$ scan points to classify, and the classification decisions must be defensible in a court of law. This means that classification is routinely performed by human readers - a monumental and time consuming process. The application of a CNN with fast computation of meaningful Shapley values provides a potential alternative to the classification. This research demonstrates the realistic, accurate and fast computation of Shapley values for this massive task

q-bio.QM

Simulating realistic short tandem repeat capillary electrophoretic signal using a generative adversarial network

DNA profiles are made up from multiple series of electrophoretic signal measuring fluorescence over time. Typically, human DNA analysts 'read' DNA profiles using their experience to distinguish instrument noise, artefactual signal, and signal corresponding to DNA fragments of interest. Recent work has developed an artificial neural network, ANN, to carry out the task of classifying fluorescence types into categories in DNA profile electrophoretic signal. But the creation of the necessarily large amount of labelled training data for the ANN is time consuming and expensive, and a limiting factor in the ability to robustly train the ANN. If realistic, prelabelled, training data could be simulated then this would remove the barrier to training an ANN with high efficacy. Here we develop a generative adversarial network, GAN, modified from the pix2pix GAN to achieve this task. With 1078 DNA profiles we train the GAN and achieve the ability to simulate DNA profile information, and then use the generator from the GAN as a 'realism filter' that applies the noise and artefact elements exhibited in typical electrophoretic signal.

cs.LG

An insightful approach to bearings-only tracking in log-polar coordinates

The choice of coordinate system in a bearings-only (BO) tracking problem influences the methods used to observe and predict the state of a moving target. Modified Polar Coordinates (MPC) and Log-Polar Coordinates (LPC) have some advantages over Cartesian coordinates. In this paper, we derive closed-form expressions for the target state prior distribution after ownship manoeuvre: the mean, covariance, and higher-order moments in LPC. We explore the use of these closed-form expressions in simulation by modifying an existing BO tracker that uses the UKF. Rather than propagating sigma points, we directly substitute current values of the mean and covariance into the time update equations at the ownship turn. This modified UKF, the CFE-UKF, performs similarly to the pure UKF, verifying the closed-form expressions. The closed-form third and fourth central moments indicate non-Gaussianity of the target state when the ownship turns. By monitoring these metrics and appropriately initialising relative range error, we can achieve a desired output mean estimated range error (MRE). The availability of these higher-order moments facilitates other extensions of the tracker not possible with a standard UKF.

physics.data-an

Optimal Proposal Particle Filters for Detecting Anomalies and Manoeuvres from Two Line Element Data

Detecting anomalous behaviour of satellites is an important goal within the broader task of space situational awareness. The Two Line Element (TLE) data published by NORAD is the only widely-available, comprehensive source of data for satellite orbits. We present here a filtering approach for detecting anomalies in satellite orbits from TLE data. Optimal proposal particle filters are deployed to track the state of the satellites' orbits. New TLEs that are unlikely given our belief of the current orbital state are designated as anomalies. The change in the orbits over time is modelled using the SGP4 model with some adaptations. A model uncertainty is derived to handle the errors in SGP4 around singularities in the orbital elements. The proposed techniques are evaluated on a set of 15 satellites for which ground truth is available and the particle filters are shown to be superior at detecting the subtle in-track and cross-track manoeuvres in the simulated dataset, as well as providing a measure of uncertainty of detections.

astro-ph.EP

Capturing functional connectomics using Riemannian partial least squares

For neurological disorders and diseases, functional and anatomical connectomes of the human brain can be used to better inform targeted interventions and treatment strategies. Functional magnetic resonance imaging (fMRI) is a non-invasive neuroimaging technique that captures spatio-temporal brain function through blood flow over time. FMRI can be used to study the functional connectome through the functional connectivity matrix; that is, Pearson's correlation matrix between time series from the regions of interest of an fMRI image. One approach to analysing functional connectivity is using partial least squares (PLS), a multivariate regression technique designed for high-dimensional predictor data. However, analysing functional connectivity with PLS ignores a key property of the functional connectivity matrix; namely, these matrices are positive definite. To account for this, we introduce a generalisation of PLS to Riemannian manifolds, called R-PLS, and apply it to symmetric positive definite matrices with the affine invariant geometry. We apply R-PLS to two functional imaging datasets: COBRE, which investigates functional differences between schizophrenic patients and healthy controls, and; ABIDE, which compares people with autism spectrum disorder and neurotypical controls. Using the variable importance in the projection statistic on the results of R-PLS, we identify key functional connections in each dataset that are well represented in the literature. Given the generality of R-PLS, this method has potential to open up new avenues for multi-model imaging analysis linking structural and functional connectomics.

stat.ME

Multivariate distance matrix regression for a manifold-valued response variable

In this paper, we propose the use of geodesic distances in conjunction with multivariate distance matrix regression, called geometric-MDMR, as a powerful first step analysis method for manifold-valued data. Manifold-valued data is appearing more frequently in the literature from analyses of earthquake to analysing brain patterns. Accounting for the structure of this data increases the complexity of your analysis, but allows for much more interpretable results in terms of the data. To test geometric-MDMR, we develop a method to simulate functional connectivity matrices for fMRI data to perform a simulation study, which shows that our method outperforms the current standards in fMRI analysis.

stat.ME

Comparison of three Statistical Classification Techniques for Maser Identification

We applied three statistical classification techniques - linear discriminant analysis (LDA), logistic regression and random forests - to three astronomical datasets associated with searches for interstellar masers. We compared the performance of these methods in identifying whether specific mid-infrared or millimetre continuum sources are likely to have associated interstellar masers. We also discuss the ease, or otherwise, with which the results of each classification technique can be interpreted. Non-parametric methods have the potential to make accurate predictions when there are complex relationships between critical parameters. We found that for the small datasets the parametric methods logistic regression and LDA performed best, for the largest dataset the non-parametric method of random forests performed with comparable accuracy to parametric techniques, rather than any significant improvement. This suggests that at least for the specific examples investigated here accuracy of the predictions obtained is not being limited by the use of parametric models. We also found that for LDA, transformation of the data to match a normal distribution in the input parameters led to big improvements in accuracy. The different classification techniques had significant overlap in their predictions, further astronomical observations will enable the accuracy of these predictions to be tested.

astro-ph.IM