SearcharxivSearch

arXiv subjects

Daniel J. Graham

Publications and source records attributed to Daniel J. Graham.

At least 19 recordsLinked to original sources

Quantifying the Causal Operational Determinants of Service Reliability in Urban Rail Transit: Evidence from Panel Double/Debiased Machine Learning

Urban rail transit reliability is a critical measure of system performance, yet its causal determinants remain poorly quantified due to high-dimensional and interdependent influencing factors. This study investigates reliability patterns across 46 international metro operators between 1994 and 2024 using the CoMET benchmarking database, incorporating more than 90 candidate variables spanning technical, operational, financial, environmental, and macroeconomic conditions. Based on domain knowledge, literature synthesis, and variable construction, four operational determinants are designed to capture three mechanisms: demand pressure, service supply, and demand-supply imbalance, while the remaining variables are screened and incorporated as confounders where theoretically appropriate. Double/Debiased Machine Learning (DML) adapted for panel data is introduced to urban rail reliability analysis to quantify the net causal effects of these determinants under complex and nonlinear relationships. The framework combines flexible machine learning with panel fixed or random effects within-operator temporal variation, reducing bias from high-dimensional confounding, model misspecification, and unobserved operator heterogeneity. The results identify three distinct operational mechanisms. Higher passenger demand intensity increases incident rates by 0.38% (p<0.001). On the supply side, greater fleet supply adequacy and car-based operational intensity reduce incident rates by 0.52% (p<0.05) and 0.80% (p<0.01), respectively. Capacity utilization, which reflects the imbalance between demand and available supply, increases incident rates by 0.49% (p<0.001). These findings show that metro reliability depends not only on the level of demand or supply alone, but also on whether service provision keeps pace with passenger demand.

stat.AP

Copy-Spread-Annihilate Dynamics in Degree-Assortative Networks

In many systems, communication proceeds by broadcasting rather than single source-target routing, but network structures that maximize signal lifetime are not well understood. Degree correlations are known to influence robustness and spreading, yet their effect on signal persistence has remained unclear. Here we introduce Copy-Spread-Annihilate dynamics, a minimal synchronous broadcasting model with annihilation. We show that signal lifetimes vary non-monotonically with assortativity and are maximized near neutral assortativity, where hub-driven amplification is strong but annihilation via short cycles is still limited. Applying this framework to the mouse connectome suggests assortativity as a structural control parameter for broadcast signal persistence in brain-like and other complex networks.

q-bio.NC

GeMA: Learning Latent Manifold Frontiers for Benchmarking Complex Systems

Benchmarking the performance of complex systems such as rail networks, renewable generation assets and national economies is central to transport planning, regulation and macroeconomic analysis. Classical frontier methods, notably Data Envelopment Analysis (DEA) and Stochastic Frontier Analysis (SFA), estimate an efficient frontier in the observed input-output space and define efficiency as distance to this frontier, but rely on restrictive assumptions on the production set and only indirectly address heterogeneity and scale effects. We propose Geometric Manifold Analysis (GeMA), a latent manifold frontier framework implemented via a productivity-manifold variational autoencoder (ProMan-VAE). Instead of specifying a frontier function in the observed space, GeMA represents the production set as the boundary of a low-dimensional manifold embedded in the joint input-output space. A split-head encoder learns latent variables that capture technological structure and operational inefficiency. Efficiency is evaluated with respect to the learned manifold, endogenous peer groups arise as clusters in latent technology space, a quotient construction supports scale-invariant benchmarking, and a local certification radius, derived from the decoder Jacobian and a Lipschitz bound, quantifies the geometric robustness of efficiency scores. We validate GeMA on synthetic data with non-convex frontiers, heterogeneous technologies and scale bias, and on four real-world case studies: global urban rail systems (COMET), British rail operators (ORR), national economies (Penn World Table) and a high-frequency wind-farm dataset. Across these domains GeMA behaves comparably to established methods when classical assumptions hold, and provides additional insight in settings with pronounced heterogeneity, non-convexity or size-related bias.

cs.LG

Interval Prediction of Annual Average Daily Traffic on Local Roads via Quantile Random Forest with High-Dimensional Spatial Data

Accurate annual average daily traffic (AADT) data are vital for transport planning and infrastructure management. However, automatic traffic detectors across national road networks often provide incomplete coverage, leading to underrepresentation of minor roads. While recent machine learning advances have improved AADT estimation at unmeasured locations, most models produce only point predictions and overlook estimation uncertainty. This study addresses that gap by introducing an interval prediction approach that explicitly quantifies predictive uncertainty. We integrate a Quantile Random Forest model with Principal Component Analysis to generate AADT prediction intervals, providing plausible traffic ranges bounded by estimated minima and maxima. Using data from over 2,000 minor roads in England and Wales, and evaluated with specialized interval metrics, the proposed method achieves an interval coverage probability of 88.22%, a normalized average width of 0.23, and a Winkler Score of 7,468.47. By combining machine learning with spatial and high-dimensional analysis, this framework enhances both the accuracy and interpretability of AADT estimation, supporting more robust and informed transport planning.

stat.ML

A longitudinal Bayesian framework for estimating causal dose-response relationships

Existing causal methods for time-varying exposure and time-varying confounding focus on estimating the average causal effect of a time-varying binary treatment on an end-of-study outcome, offering limited tools for characterizing marginal causal dose-response relationships under continuous exposures. We propose a scalable, nonparametric Bayesian framework for estimating marginal longitudinal causal dose-response functions with repeated outcome measurements. Our approach targets the average potential outcome at any fixed dose level and accommodates time-varying confounding through the generalized propensity score. The proposed approach embeds a Dirichlet process specification within a generalized estimating equations structure, capturing temporal correlation while making minimal assumptions about the functional form of the continuous exposure. We apply the proposed methods to monthly metro ridership and COVID-19 case data from major international cities, identifying causal relationships and the dose-response patterns between higher ridership and increased case counts.

stat.ME

Message interaction dynamics covary with brain volume across mammalian connectomes

Brain network communication models typically assume that signals propagate independently, despite the high network density and small diameter of mammalian connectomes, where interactions among simultaneously propagating messages are likely. We investigate these interactions using the copy-spread-annihilate (CSA) model, a synchronous Markovian message-passing process in which messages spread through the binarized network while undergoing collision-dependent deletion. Simulations on a large comparative dataset of mammalian connectomes show that CSA dynamics produce robust positively-skewed lognormal distributions of message survival across species. Despite using only binary network topology without spatial embedding, message survival accounts for over half of the variance in brain volume across mammal species and outperforms a broad range of established graph-theoretic measures and an alternative communication model with interacting messages. Degree-preserving switch randomization weakens this relationship, indicating that higher-order structural organization contributes to the observed scaling law. We describe a mechanistic explanation of model dynamics that suggests message creation differences among networks shapes message survival. Together, these findings suggest that interactions among propagating messages constitute an important and underexplored determinant of communication dynamics in mammalian brain networks.

q-bio.NC

Diamine Surface Passivation and Post-Annealing Enhance Performance of Silicon-Perovskite Tandem Solar Cells

We show that the use of 1,3-diaminopropane (DAP) as a chemical modifier at the perovskite/electron-transport layer (ETL) interface enhances the power conversion efficiency (PCE) of 1.7 eV bandgap FACs mixed-halide perovskite single-junction cells, primarily by boosting the open-circuit voltage (VOC) from 1.06 V to 1.15 V. Adding a post-processing annealing step after C60 evaporation, further improves the fill factor (FF) by 20% from the control to the DAP + post-annealing devices. Using hyperspectral photoluminescence microscopy, we demonstrate that annealing helps improve compositional homogeneity at the top and bottom interfaces of the solar cell, which prevents detrimental bandgap pinning in the devices and improves C60 adhesion. Using time-of-flight secondary ion mass spectrometry, we show that DAP reacts with formamidinium present near the surface of the perovskite lattice to form a larger molecular cation, 1,4,5,6-tetrahydropyrimidinium (THP) that remains at the interface. Combining the use of DAP and the annealing of C60 interface, we fabricate Si-perovskite tandems with PCE of 25.29%, compared to 23.26% for control devices. Our study underscores the critical role of chemical reactivity and thermal post-processing of the C60/Lewis-base passivator interface in minimizing device losses and advancing solar-cell performance of wide-bandgap mixed-cation mixed-halide perovskite for tandem application.

cond-mat.mtrl-sci

MSCT: Addressing Time-Varying Confounding with Marginal Structural Causal Transformer for Counterfactual Post-Crash Traffic Prediction

Traffic crashes profoundly impede traffic efficiency and pose economic challenges. Accurate prediction of post-crash traffic status provides essential information for evaluating traffic perturbations and developing effective solutions. Previous studies have established a series of deep learning models to predict post-crash traffic conditions, however, these correlation-based methods cannot accommodate the biases caused by time-varying confounders and the heterogeneous effects of crashes. The post-crash traffic prediction model needs to estimate the counterfactual traffic speed response to hypothetical crashes under various conditions, which demonstrates the necessity of understanding the causal relationship between traffic factors. Therefore, this paper presents the Marginal Structural Causal Transformer (MSCT), a novel deep learning model designed for counterfactual post-crash traffic prediction. To address the issue of time-varying confounding bias, MSCT incorporates a structure inspired by Marginal Structural Models and introduces a balanced loss function to facilitate learning of invariant causal features. The proposed model is treatment-aware, with a specific focus on comprehending and predicting traffic speed under hypothetical crash intervention strategies. In the absence of ground-truth data, a synthetic data generation procedure is proposed to emulate the causal mechanism between traffic speed, crashes, and covariates. The model is validated using both synthetic and real-world data, demonstrating that MSCT outperforms state-of-the-art models in multi-step-ahead prediction performance. This study also systematically analyzes the impact of time-varying confounding bias and dataset distribution on model performance, contributing valuable insights into counterfactual prediction for intelligent transportation systems.

cs.LG

Causal resilience curves: A data-driven framework for quantifying the spatiotemporal impacts of metro service disruptions

Urban metro systems move vast numbers of passengers with a high level of efficiency in resource use, but frequently experience disruptions that result in delays, crowding, and deterioration in passenger satisfaction and patronage. To quantify these adverse consequences, this paper presents a novel, data-driven causal inference framework to measure metro resilience by estimating both the direct and spillover effects of service disruptions on passenger demand, journey time, travel speed and on-board crowding. By integrating high-frequency smart card data into a synthetic control design, we use weighted non-disrupted days to construct unbiased counterfactuals, which resolves confounding factors and accurately captures disruption propagation across the network. The impact estimates are further translated into station-level causal resilience curves that reveal spatial heterogeneity in the temporal patterns of degradation and recovery across locations, providing metro operators with actionable insights for targeted interventions and resource allocation. A case study of the Hong Kong MTR demonstrates the framework's superiority over naive typical-day comparisons and machine-learning benchmarks in delivering unbiased resilience curves. This paper is the first to derive causal estimates of dynamic metro resilience. This practical tool can be generalised to evaluate resilience in a broad range of public transport systems.

stat.AP

Reduced Recombination via Tunable Surface Fields in Perovskite Solar Cells

The ability to reduce energy loss at semiconductor surfaces through passivation or surface field engineering has become an essential step in the manufacturing of efficient photovoltaic (PV) and optoelectronic devices. Similarly, surface modification of emerging halide perovskites with quasi-2D heterostructures is now ubiquitous to achieve PV power conversion efficiencies (PCEs) > 22% and has enabled single-junction PV devices to reach 25.7%, yet a fundamental understanding to how these treatments function is still generally lacking. This has established a bottleneck for maximizing beneficial improvements as no concrete selection and design rules currently exist. Here we uncover a new type of tunable passivation strategy and mechanism found in perovskite PV devices that were the first to reach the > 25% PCE milestone, which is enabled by surface treating a bulk perovskite layer with hexylammonium bromide (HABr). We uncover the simultaneous formation of an iodide-rich 2D layer along with a Br halide gradient achieved through partial halide exchange that extends from defective surfaces and grain boundaries into the bulk layer. We demonstrate and directly visualize the tunability of both the 2D layer thickness, halide gradient, and band structure using a unique combination of depth-sensitive nanoscale characterization techniques. We show that the optimization of this interface can extend the charge carrier lifetime to values > 30 μs, which is the longest value reported for a direct bandgap semiconductor (GaAs, InP, CdTe) over the past 50 years. Importantly, this work reveals an entirely new strategy and knob for optimizing and tuning recombination and charge transport at semiconductor interfaces and will likely establish new frontiers in achieving the next set of perovskite device performance records.

physics.app-ph

Optimal congestion control strategies for near-capacity urban metros: informing intervention via fundamental diagrams

Congestion; operational delays due to a vicious circle of passenger-congestion and train-queuing; is an escalating problem for metro systems because it has negative consequences from passenger discomfort to eventual mode-shifts. Congestion arises due to large volumes of passenger boardings and alightings at bottleneck stations, which may lead to increased stopping times at stations and consequent queuing of trains upstream, further reducing line throughput and implying an even greater accumulation of passengers at stations. Alleviating congestion requires control strategies such as regulating the inflow of passengers entering bottleneck stations. The availability of large-scale smartcard and train movement data from day-to-day operations facilitates the development of models that can inform such strategies in a data-driven way. In this paper, we propose to model station-level passenger-congestion via empirical passenger boarding-alightings and train flow relationships, henceforth, fundamental diagrams (FDs). We emphasise that estimating FDs using station-level data is empirically challenging due to confounding biases arising from the interdependence of operations at different stations, which obscures the true sources of congestion in the network. We thus adopt a causal statistical modelling approach to produce FDs that are robust to confounding and as such suitable to properly inform control strategies. The closest antecedent to the proposed model is the FD for road traffic networks, which informs traffic management strategies, for instance, via locating the optimum operation point. Our analysis of data from the Mass Transit Railway, Hong Kong indicates the existence of concave FDs at identified bottleneck stations, and an associated critical level of boarding-alightings above which congestion sets-in unless there is an intervention.

stat.AP

Willingness to Pay and Attitudinal Preferences of Indian Consumers for Electric Vehicles

Consumer preference elicitation is critical to devise effective policies for the diffusion of electric vehicles (EVs) in India. This study contributes to the EV demand literature in the Indian context by (a) analysing the EV attributes and attitudinal factors of Indian car buyers that determine consumers' preferences for EVs, (b) estimating Indian consumers' willingness to pay (WTP) to buy EVs with improved attributes, and c) quantifying how the reference dependence affects the WTP estimates. We adopt a hybrid choice modelling approach for the above analysis. The results indicate that accounting for reference dependence provides more realistic WTP estimates than the standard utility estimation approach. Our results suggest that Indian consumers are willing to pay an additional USD 10-34 in the purchase price to reduce the fast charging time by 1 minute, USD 7-40 to add a kilometre to the driving range of EVs at 200 kilometres, and USD 104-692 to save USD 1 per 100 kilometres in operating cost. These estimates and the effect of attitudes on the likelihood to adopt EVs provide insights about EV design, marketing strategies, and pro-EV policies (e.g., specialised lanes and reserved parking for EVs) to expedite the adoption of EVs in India.

econ.GN

Revisiting the empirical fundamental relationship of traffic flow for highways using a causal econometric approach

The fundamental relationship of traffic flow is empirically estimated by fitting a regression curve to a cloud of observations of traffic variables. Such estimates, however, may suffer from the confounding/endogeneity bias due to omitted variables such as driving behaviour and weather. To this end, this paper adopts a causal approach to obtain an unbiased estimate of the fundamental flow-density relationship using traffic detector data. In particular, we apply a Bayesian non-parametric spline-based regression approach with instrumental variables to adjust for the aforementioned confounding bias. The proposed approach is benchmarked against standard curve-fitting methods in estimating the flow-density relationship for three highway bottlenecks in the United States. Our empirical results suggest that the saturated (or hypercongested) regime of the estimated flow-density relationship using correlational curve fitting methods may be severely biased, which in turn leads to biased estimates of important traffic control inputs such as capacity and capacity-drop. We emphasise that our causal approach is based on the physical laws of vehicle movement in a traffic stream as opposed to a demand-supply framework adopted in the economics literature. By doing so, we also aim to conciliate the engineering and economics approaches to this empirical problem. Our results, thus, have important implications both for traffic engineers and transport economists.

econ.EM

Non-parametric Bayesian inference via loss functions under model misspecification

In the usual Bayesian setting, a full probabilistic model is required to link the data and parameters, and the form of this model and the inference and prediction mechanisms are specified via de Finetti's representation. In general, such a formulation is not robust to model misspecification of its component parts. An alternative approach is to draw inference based on loss functions, where the quantity of interest is defined as a minimizer of some expected loss, and to construct posterior distributions based on the loss-based formulation; this strategy underpins the construction of the Gibbs posterior. We develop a Bayesian non-parametric approach; specifically, we generalize the Bayesian bootstrap, and specify a Dirichlet process model for the distribution of the observables. We implement this using direct prior-to-posterior calculations, but also using predictive sampling. We also study the assessment of posterior validity for non-standard Bayesian calculations. We show that the developed non-standard Bayesian updating procedures yield valid posterior distributions in terms of consistency and asymptotic normality under model misspecification. Simulation studies show that the proposed methods can recover the true value of the parameter under misspecification.

stat.ME

Analysing the causal effect of London cycle superhighways on traffic congestion

Transport operators have a range of intervention options available to improve or enhance their networks. Such interventions are often made in the absence of sound evidence on resulting outcomes. Cycling superhighways were promoted as a sustainable and healthy travel mode, one of the aims of which was to reduce traffic congestion. Estimating the impacts that cycle superhighways have on congestion is complicated due to the non-random assignment of such intervention over the transport network. In this paper, we analyse the causal effect of cycle superhighways utilising pre-intervention and post-intervention information on traffic and road characteristics along with socio-economic factors. We propose a modeling framework based on the propensity score and outcome regression model. The method is also extended to the doubly robust set-up. Simulation results show the superiority of the performance of the proposed method over existing competitors. The method is applied to analyse a real dataset on the London transport network. The methodology proposed can assist in effective decision making to improve network performance.

stat.AP

A Causal Inference Approach to Measure the Vulnerability of Urban Metro Systems

Transit operators need vulnerability measures to understand the level of service degradation under disruptions. This paper contributes to the literature with a novel causal inference approach for estimating station-level vulnerability in metro systems. The empirical analysis is based on large-scale data on historical incidents and population-level passenger demand. This analysis thus obviates the need for assumptions made by previous studies on human behaviour and disruption scenarios. We develop four empirical vulnerability metrics based on the causal impact of disruptions on travel demand, average travel speed and passenger flow distribution. Specifically, the proposed metrics based on the irregularity in passenger flow distribution extends the scope of vulnerability measurement to the entire trip distribution, instead of just analysing the disruption impact on the entry or exit demand (that is, moments of the trip distribution). The unbiased estimates of disruption impact are obtained by adopting a propensity score matching method, which adjusts for the confounding biases caused by non-random occurrence of disruptions. An application of the proposed framework to the London Underground indicates that the vulnerability of a metro station depends on the location, topology, and other characteristics. We find that, in 2013, central London stations are more vulnerable in terms of travel demand loss. However, the loss of average travel speed and irregularity in relative passenger flows reveal that passengers from outer London stations suffer from longer individual delays due to lack of alternative routes.

stat.AP

Fast Bayesian Estimation of Spatial Count Data Models

Spatial count data models are used to explain and predict the frequency of phenomena such as traffic accidents in geographically distinct entities such as census tracts or road segments. These models are typically estimated using Bayesian Markov chain Monte Carlo (MCMC) simulation methods, which, however, are computationally expensive and do not scale well to large datasets. Variational Bayes (VB), a method from machine learning, addresses the shortcomings of MCMC by casting Bayesian estimation as an optimisation problem instead of a simulation problem. Considering all these advantages of VB, a VB method is derived for posterior inference in negative binomial models with unobserved parameter heterogeneity and spatial dependence. Pólya-Gamma augmentation is used to deal with the non-conjugacy of the negative binomial likelihood and an integrated non-factorised specification of the variational distribution is adopted to capture posterior dependencies. The benefits of the proposed approach are demonstrated in a Monte Carlo study and an empirical application on estimating youth pedestrian injury counts in census tracts of New York City. The VB approach is around 45 to 50 times faster than MCMC on a regular eight-core processor in a simulation and an empirical study, while offering similar estimation and predictive accuracy. Conditional on the availability of computational resources, the embarrassingly parallel architecture of the proposed VB method can be exploited to further accelerate its estimation by up to 20 times.

stat.ME

A Dynamic Choice Model with Heterogeneous Decision Rules: Application in Estimating the User Cost of Rail Crowding

Crowding valuation of subway riders is an important input to various supply-side decisions of transit operators. The crowding cost perceived by a transit rider is generally estimated by capturing the trade-off that the rider makes between crowding and travel time while choosing a route. However, existing studies rely on static compensatory choice models and fail to account for inertia and the learning behaviour of riders. To address these challenges, we propose a new dynamic latent class model (DLCM) which (i) assigns riders to latent compensatory and inertia/habit classes based on different decision rules, (ii) enables transitions between these classes over time, and (iii) adopts instance-based learning theory to account for the learning behaviour of riders. We use the expectation-maximisation algorithm to estimate DLCM, and the most probable sequence of latent classes for each rider is retrieved using the Viterbi algorithm. The proposed DLCM can be applied in any choice context to capture the dynamics of decision rules used by a decision-maker. We demonstrate its practical advantages in estimating the crowding valuation of an Asian metro's riders. To calibrate the model, we recover the daily route preferences and in-vehicle crowding experiences of regular metro riders using a two-month-long smart card and vehicle location data. The results indicate that the average rider follows the compensatory rule on only 25.5% of route choice occasions. DLCM estimates also show an increase of 47% in metro riders' valuation of travel time under extremely crowded conditions relative to that under uncrowded conditions.

stat.AP