SearcharxivSearch

arXiv subjects

Ben Moews

Publications and source records attributed to Ben Moews.

At least 19 recordsLinked to original sources

Data-Driven Measures of High-Frequency Trading

Public data do not identify high-frequency trading (HFT), and standard proxies do not separate liquidity-supplying from liquidity-demanding strategies. We overcome this measurement challenge by training machine learning models on proprietary Nasdaq data to map observed HFT activity to public intraday variables. Applying this mapping, we generate daily measures of liquidity-supplying and liquidity-demanding HFT for all U.S. stocks from 2010 to 2023. The measures largely subsume standard proxies and capture time-series variation that those proxies miss. Using proprietary Euronext Paris data, we provide evidence that the approach generalizes across markets and remains predictive years after training. The 14-year panel lets us study HFT and market quality over time. Supply-side HFT is consistently associated with greater pre-announcement information acquisition, more informed trading, and lower bid-ask spreads, while demand-side HFT is associated with the opposite patterns. During COVID-19, HFT-supplied liquidity remained resilient and its association with lower spreads strengthened.

q-fin.GN

Crime reduction through public healthcare: Interpretable machine learning for mental health service impacts in Greater London

The relationship between crime, mental health service access, and socioeconomic deprivation in publicly-funded healthcare systems allowing impactful policy interventions offers an alternative lens to crime prevention that remains underexplored. We address this critical gap through an analysis of street-level crime data, mental health referral information, and socioeconomic metrics across Greater London, using both traditional statistical methods and machine learning techniques to identify relevant relationships and spatial patterns to reveal a persistent positive association between crime rates and mental health referrals as a proxy for service access. The prevailing prevention hypothesis is contrasted with a nuanced U-shaped relationship suggesting a contrast between preventive effects at lower service levels and demand-driven responses to crime exposure for higher referral rates. Subsequent analyses, focussing on explainable artificial intelligence, show distinct crime category patterns, with a cluster analysis identifying four borough typologies with distinct combinations of crime rates, mental health service access, and deprivation levels, requiring multifaceted approaches rather than universal solutions. This research provides one of the first comprehensive studies on this topic for the UK's publicly-funded healthcare system and introduces interpretation-oriented approaches to uncover the patterns essential to evidence-based policies.

stat.AP

Weak Lensing by Photometric Density Ridges

Ridges in galaxy density fields measured by photometric surveys are 2D projections of filaments in the cosmic web, and so should lens light from background galaxies. We report on a detection of this effect in Dark Energy Survey Year 3 data at high significance, though not independently of galaxy-galaxy lensing. We describe improvements to the existing subspace-constrained mean shift algorithm to locate these ridges efficiently at scale, and examine the dependence of the signal in simulations on cosmological and algorithmic parameters. We find that it depends primarily on $S_8=\sigma_8 \left( \Omega_m / 0.3 \right)^{1/2}$, and discuss improvements to our methodology that would be needed to allow precision parameter estimation.

astro-ph.CO

Permanent and transitory crime risk in variable-density hot spot analysis

Crime prevention measures, aiming for the effective and efficient spending of public resources, rely on the empirical analysis of spatial and temporal data for public safety outcomes. We perform a variable-density cluster analysis on crime incident reports in the City of Chicago for the years 2001--2022 to investigate changes in crime share composition for hot spots of different densities. Contributing to and going beyond the existing wealth of research on criminological applications in the operational research literature, we study the evolution of crime type shares in clusters over the course of two decades and demonstrate particularly notable impacts of the COVID-19 pandemic and its associated social contact avoidance measures, as well as a dependence of these effects on the primary function of city areas. Our results also indicate differences in the relative difficulty to address specific crime types, and an analysis of spatial autocorrelations further shows variations in incident uniformity between clusters and outlier areas at different distance radii. We discuss our findings in the context of the interplay between operational research and criminal justice, the practice of hot spot policing and public safety optimization, and the factors contributing to, and challenges and risks due to, data biases as an often neglected factor in criminological applications.

stat.AP

Public transport challenges and technology-assisted accessibility for visually impaired elderly residents in urban environments

Independent navigation is central to social participation and health for vulnerable populations. While historic cities such as Edinburgh often feature well-established public transport systems, urban accessibility challenges remain and are exacerbated by complex landscapes, especially for groups with multiple vulnerabilities such as the visually impaired elderly. With limited research examining how real-time data feeds and artificial intelligence in this context, we address this gap through a mixed-methods approach. Our spatio-temporal analyses make use of statistical and machine learning techniques to investigate network coverage, service patterns, and density profiles through live-recorded data. This is combined with a qualitative thematic analysis of semi-structured interviews with the target group, as well as links to spatial cognition theory. The results demonstrate the highly centralised nature of the city's transport system, the significance of memory-based navigation, and the lack of travel information in usable formats. We also find that participants already use navigation technology to varying degrees and express a willingness to adopt artificial intelligence. Our findings highlight the importance of dynamic tools to meaningfully improve independent travel, as well as limitations due to the recurring problem of specific accessibility data, for example for facilities, often not being collected and stored.

cs.HC

Evaluating utility in synthetic banking microdata applications

Financial regulators such as central banks collect vast amounts of data, but access to the resulting fine-grained banking microdata is severely restricted by banking secrecy laws. Recent developments have resulted in mechanisms that generate faithful synthetic data, but current evaluation frameworks lack a focus on the specific challenges of banking institutions and microdata. We develop a framework that considers the utility and privacy requirements of regulators, and apply this to financial usage indices, term deposit yield curves, and credit card transition matrices. Using the Central Bank of Paraguay's data, we provide the first implementation of synthetic banking microdata using a central bank's collected information, with the resulting synthetic datasets for all three domain applications being publicly available and featuring information not yet released in statistical disclosure. We find that applications less susceptible to post-processing information loss, which are based on frequency tables, are particularly suited for this approach, and that marginal-based inference mechanisms to outperform generative adversarial network models for these applications. Our results demonstrate that synthetic data generation is a promising privacy-enhancing technology for financial regulators seeking to complement their statistical disclosure, while highlighting the crucial role of evaluating such endeavors in terms of utility and privacy requirements.

q-fin.CP

SCADDA: Spatio-temporal cluster analysis with density-based distance augmentation and its application to fire carbon emissions

Spatio-temporal clustering occupies an established role in various fields dealing with geospatial analysis, spanning from healthcare analysis to environmental science. One major challenge are applications in which cluster assignments are dependent on local densities, meaning that higher-density areas should be treated more strictly for spatial clustering and vice versa. Meeting this need, we describe and implement an extended method that covers continuous and adaptive distance rescaling based on kernel density estimates and the orthodromic metric, as well as the distance between time series via dynamic time warping. In doing so, we provide the wider research community, as well as practitioners, with a novel approach to solve an existing challenge as well as an easy-to-handle and robust open-source software tool. The resulting implementation is highly customizable to suit different application cases, and we verify and test the latter on both an idealized scenario and the recreation of prior work on broadband antibiotics prescriptions in Scotland to demonstrate well-behaved comparative performance. Following this, we apply our approach to fire emissions in Sub-Saharan Africa using data from Earth-observing satellites, and show our implementation's ability to uncover seasonality shifts in carbon emissions of subgroups as a result of time series-driven cluster splits.

stat.CO

Physics-informed neural networks in the recreation of hydrodynamic simulations from dark matter

Physics-informed neural networks have emerged as a coherent framework for building predictive models that combine statistical patterns with domain knowledge. The underlying notion is to enrich the optimization loss function with known relationships to constrain the space of possible solutions. Hydrodynamic simulations are a core constituent of modern cosmology, while the required computations are both expensive and time-consuming. At the same time, the comparatively fast simulation of dark matter requires fewer resources, which has led to the emergence of machine learning algorithms for baryon inpainting as an active area of research; here, recreating the scatter found in hydrodynamic simulations is an ongoing challenge. This paper presents the first application of physics-informed neural networks to baryon inpainting by combining advances in neural network architectures with physical constraints, injecting theory on baryon conversion efficiency into the model loss function. We also introduce a punitive prediction comparison based on the Kullback-Leibler divergence, which enforces scatter reproduction. By simultaneously extracting the complete set of baryonic properties for the Simba suite of cosmological simulations, our results demonstrate improved accuracy of baryonic predictions based on dark matter halo properties, successful recovery of the fundamental metallicity relation, and retrieve scatter that traces the target simulation's distribution.

astro-ph.CO

On random number generators and practical market efficiency

Modern mainstream financial theory is underpinned by the efficient market hypothesis, which posits the rapid incorporation of relevant information into asset pricing. Limited prior studies in the operational research literature have investigated tests designed for random number generators to check for these informational efficiencies. Treating binary daily returns as a hardware random number generator analogue, tests of overlapping permutations have indicated that these time series feature idiosyncratic recurrent patterns. Contrary to prior studies, we split our analysis into two streams at the annual and company level, and investigate longer-term efficiency over a larger time frame for Nasdaq-listed public companies to diminish the effects of trading noise and allow the market to realistically digest new information. Our results demonstrate that information efficiency varies across years and reflects large-scale market impacts such as financial crises. We also show the proximity to results of a well-tested pseudo-random number generator, discuss the distinction between theoretical and practical market efficiency, and find that the statistical qualification of stock-separated returns in support of the efficient market hypothesis is dependent on the driving factor of small inefficient subsets that skew market assessments.

stat.AP

Photometric Redshift Uncertainties in Weak Gravitational Lensing Shear Analysis: Models and Marginalization

Recovering credible cosmological parameter constraints in a weak lensing shear analysis requires an accurate model that can be used to marginalize over nuisance parameters describing potential sources of systematic uncertainty, such as the uncertainties on the sample redshift distribution $n(z)$. Due to the challenge of running Markov Chain Monte-Carlo (MCMC) in the high dimensional parameter spaces in which the $n(z)$ uncertainties may be parameterized, it is common practice to simplify the $n(z)$ parameterization or combine MCMC chains that each have a fixed $n(z)$ resampled from the $n(z)$ uncertainties. In this work, we propose a statistically-principled Bayesian resampling approach for marginalizing over the $n(z)$ uncertainty using multiple MCMC chains. We self-consistently compare the new method to existing ones from the literature in the context of a forecasted cosmic shear analysis for the HSC three-year shape catalog, and find that these methods recover similar cosmological parameter constraints, implying that using the most computationally efficient of the approaches is appropriate. However, we find that for datasets with the constraining power of the full HSC survey dataset (and, by implication, those upcoming surveys with even tighter constraints), the choice of method for marginalizing over $n(z)$ uncertainty among the several methods from the literature may significantly impact the statistical uncertainties on cosmological parameters, and a careful model selection is needed to ensure credible parameter intervals.

astro-ph.CO

Filaments of crime: Informing policing via thresholded ridge estimation

Objectives: We introduce a new method for reducing crime in hot spots and across cities through ridge estimation. In doing so, our goal is to explore the application of density ridges to hot spots and patrol optimization, and to contribute to the policing literature in police patrolling and crime reduction strategies. Methods: We make use of the subspace-constrained mean shift algorithm, a recently introduced approach for ridge estimation further developed in cosmology, which we modify and extend for geospatial datasets and hot spot analysis. Our experiments extract density ridges of Part I crime incidents from the City of Chicago during the year 2018 and early 2019 to demonstrate the application to current data. Results: Our results demonstrate nonlinear mode-following ridges in agreement with broader kernel density estimates. Using early 2019 incidents with predictive ridges extracted from 2018 data, we create multi-run confidence intervals and show that our patrol templates cover around 94% of incidents for 0.1-mile envelopes around ridges, quickly rising to near-complete coverage. We also develop and provide researchers, as well as practitioners, with a user-friendly and open-source software for fast geospatial density ridge estimation. Conclusions: We show that ridges following crime report densities can be used to enhance patrolling capabilities. Our empirical tests show the stability of ridges based on past data, offering an accessible way of identifying routes within hot spots instead of patrolling epicenters. We suggest further research into the application and efficacy of density ridges for patrolling.

stat.AP

Ridges in the Dark Energy Survey for cosmic trough identification

Cosmic voids and their corresponding redshift-projected mass densities, known as troughs, play an important role in our attempt to model the large-scale structure of the Universe. Understanding these structures enables us to compare the standard model with alternative cosmologies, constrain the dark energy equation of state, and distinguish between different gravitational theories. In this paper, we extend the subspace-constrained mean shift algorithm, a recently introduced method to estimate density ridges, and apply it to 2D weak lensing mass density maps from the Dark Energy Survey Y1 data release to identify curvilinear filamentary structures. We compare the obtained ridges with previous approaches to extract trough structure in the same data, and apply curvelets as an alternative wavelet-based method to constrain densities. We then invoke the Wasserstein distance between noisy and noiseless simulations to validate the denoising capabilities of our method. Our results demonstrate the viability of ridge estimation as a precursor for denoising weak lensing observables to recover the large-scale structure, paving the way for a more versatile and effective search for troughs.

astro-ph.CO

Hybrid analytic and machine-learned baryonic property insertion into galactic dark matter haloes

While cosmological dark matter-only simulations relying solely on gravitational effects are comparably fast to compute, baryonic properties in simulated galaxies require complex hydrodynamic simulations that are computationally costly to run. We explore the merging of an extended version of the equilibrium model, an analytic formalism describing the evolution of the stellar, gas, and metal content of galaxies, into a machine learning framework. In doing so, we are able to recover more properties than the analytic formalism alone can provide, creating a high-speed hydrodynamic simulation emulator that populates galactic dark matter haloes in N-body simulations with baryonic properties. While there exists a trade-off between the reached accuracy and the speed advantage this approach offers, our results outperform an approach using only machine learning for a subset of baryonic properties. We demonstrate that this novel hybrid system enables the fast completion of dark matter-only information by mimicking the properties of a full hydrodynamic suite to a reasonable degree, and discuss the advantages and disadvantages of hybrid versus machine learning-only frameworks. In doing so, we offer an acceleration of commonly deployed simulations in cosmology.

astro-ph.GA

Gaussbock: Fast parallel-iterative cosmological parameter estimation with Bayesian nonparametrics

We present and apply Gaussbock, a new embarrassingly parallel iterative algorithm for cosmological parameter estimation designed for an era of cheap parallel computing resources. Gaussbock uses Bayesian nonparametrics and truncated importance sampling to accurately draw samples from posterior distributions with an orders-of-magnitude speed-up in wall time over alternative methods. Contemporary problems in this area often suffer from both increased computational costs due to high-dimensional parameter spaces and consequent excessive time requirements, as well as the need for fine tuning of proposal distributions or sampling parameters. Gaussbock is designed specifically with these issues in mind. We explore and validate the performance and convergence of the algorithm on a fast approximation to the Dark Energy Survey Year 1 (DES Y1) posterior, finding reasonable scaling behavior with the number of parameters. We then test on the full DES Y1 posterior using large-scale supercomputing facilities, and recover reasonable agreement with previous chains, although the algorithm can underestimate the tails of poorly-constrained parameters. Additionally, we discuss and demonstrate how Gaussbock recovers complex posterior shapes very well at lower dimensions, but faces challenges to perform well on such distributions in higher dimensions. In addition, we provide the community with a user-friendly software tool for accelerated cosmological parameter estimation based on the methodology described in this paper.

astro-ph.CO

Predictive intraday correlations in stable and volatile market environments: Evidence from deep learning

Standard methods and theories in finance can be ill-equipped to capture highly non-linear interactions in financial prediction problems based on large-scale datasets, with deep learning offering a way to gain insights into correlations in markets as complex systems. In this paper, we apply deep learning to econometrically constructed gradients to learn and exploit lagged correlations among S&P 500 stocks to compare model behaviour in stable and volatile market environments, and under the exclusion of target stock information for predictions. In order to measure the effect of time horizons, we predict intraday and daily stock price movements in varying interval lengths and gauge the complexity of the problem at hand with a modification of our model architecture. Our findings show that accuracies, while remaining significant and demonstrating the exploitability of lagged correlations in stock markets, decrease with shorter prediction horizons. We discuss implications for modern finance theory and our work's applicability as an investigative tool for portfolio managers. Lastly, we show that our model's performance is consistent in volatile markets by exposing it to the environment of the recent financial crisis of 2007/2008.

q-fin.CP

On the road to percent accuracy II: calibration of the non-linear matter power spectrum for arbitrary cosmologies

We introduce an emulator approach to predict the non-linear matter power spectrum for broad classes of beyond-$Λ$CDM cosmologies, using only a suite of $Λ$CDM $N$-body simulations. By including a range of suitably modified initial conditions in the simulations, and rescaling the resulting emulator predictions with analytical `halo model reactions', accurate non-linear matter power spectra for general extensions to the standard $Λ$CDM model can be calculated. We optimise the emulator design by substituting the simulation suite with non-linear predictions from the standard {\sc halofit} tool. We review the performance of the emulator for artificially generated departures from the standard cosmology as well as for theoretically motivated models, such as $f (R)$ gravity and massive neutrinos. For the majority of cosmologies we have tested, the emulator can reproduce the matter power spectrum with errors $\lesssim 1\%$ deep into the highly non-linear regime. This work demonstrates that with a well-designed suite of $Λ$CDM simulations, extensions to the standard cosmological model can be tested in the non-linear regime without any reliance on expensive beyond-$Λ$CDM simulations.

astro-ph.CO

Stress testing the dark energy equation of state imprint on supernova data

This work determines the degree to which a standard Lambda-CDM analysis based on type Ia supernovae can identify deviations from a cosmological constant in the form of a redshift-dependent dark energy equation of state w(z). We introduce and apply a novel random curve generator to simulate instances of w(z) from constraint families with increasing distinction from a cosmological constant. After producing a series of mock catalogs of binned type Ia supernovae corresponding to each w(z) curve, we perform a standard Lambda-CDM analysis to estimate the corresponding posterior densities of the absolute magnitude of type Ia supernovae, the present-day matter density, and the equation of state parameter. Using the Kullback-Leibler divergence between posterior densities as a difference measure, we demonstrate that a standard type Ia supernova cosmology analysis has limited sensitivity to extensive redshift dependencies of the dark energy equation of state. In addition, we report that larger redshift-dependent departures from a cosmological constant do not necessarily manifest easier-detectable incompatibilities with the Lambda-CDM model. Our results suggest that physics beyond the standard model may simply be hidden in plain sight.

astro-ph.CO

Photometry of high-redshift blended galaxies using deep learning

The new generation of deep photometric surveys requires unprecedentedly precise shape and photometry measurements of billions of galaxies to achieve their main science goals. At such depths, one major limiting factor is the blending of galaxies due to line-of-sight projection, with an expected fraction of blended galaxies of up to 50%. Current deblending approaches are in most cases either too slow or not accurate enough to reach the level of requirements. This work explores the use of deep neural networks to estimate the photometry of blended pairs of galaxies in monochrome space images, similar to the ones that will be delivered by the Euclid space telescope. Using a clean sample of isolated galaxies from the CANDELS survey, we artificially blend them and train two different network models to recover the photometry of the two galaxies. We show that our approach can recover the original photometry of the galaxies before being blended with $\sim$7% accuracy without any human intervention and without any assumption on the galaxy shape. This represents an improvement of at least a factor of 4 compared to the classical SExtractor approach. We also show that forcing the network to simultaneously estimate a binary segmentation map results in a slightly improved photometry. All data products and codes will be made public to ease the comparison with other approaches on a common data set.

astro-ph.GA