SearcharxivSearch

arXiv subjects

Paul Smith

Publications and source records attributed to Paul Smith.

At least 19 recordsLinked to original sources

DMCD: Semantic-Statistical Framework for Causal Discovery

We present DMCD (DataMap Causal Discovery), a two-phase causal discovery framework that integrates LLM-based semantic drafting from variable metadata with statistical validation on observational data. In Phase I, a large language model proposes a sparse draft DAG, serving as a semantically informed prior over the space of possible causal structures. In Phase II, this draft is audited and refined via conditional independence testing, with detected discrepancies guiding targeted edge revisions. We evaluate our approach on three metadata-rich real-world benchmarks spanning industrial engineering, environmental monitoring, and IT systems analysis. Across these datasets, DMCD achieves competitive or leading performance against diverse causal discovery baselines, with particularly large gains in recall and F1 score. Probing and ablation experiments suggest that these improvements arise from semantic reasoning over metadata rather than memorization of benchmark graphs. Overall, our results demonstrate that combining semantic priors with principled statistical verification yields a high-performing and practically effective approach to causal structure learning.

cs.AI

Beyond Accuracy: A Stability-Aware Metric for Multi-Horizon Forecasting

Traditional time series forecasting methods optimize for accuracy alone. This objective neglects temporal consistency, in other words, how consistently a model predicts the same future event as the forecast origin changes. We introduce the forecast accuracy and coherence score (forecast AC score for short) for measuring the quality of probabilistic multi-horizon forecasts in a way that accounts for both multi-horizon accuracy and stability. Our score additionally allows user-specified weights to balance accuracy and consistency requirements. As an example application, we implement the score as a differentiable objective function for training seasonal auto-regressive integrated models and evaluate it on the M4 Hourly benchmark dataset. Results demonstrate consistent improvements over traditional maximum likelihood estimation. Regarding stability, the AC-optimized model generated out-of-sample forecasts with 15.8\% reduced variance over forecasts targeting the same timestamp. In terms of accuracy, the AC-optimized model achieved considerable improvements for medium-to-long-horizon forecasts. While one-step-ahead forecasts exhibited a 3.9\% increase in MSE, forecasts from horizon three onward experienced improved accuracy, with a peak improvement of approximately 6\% in MSE at horizons 9-12. These results indicate that our metric successfully trains models to produce more stable and accurate multi-step forecasts in exchange for a relatively small degradation in one-step-ahead performance.

cs.LG

Causify DataFlow: A Framework For High-performance Machine Learning Stream Computing

We present DataFlow, a computational framework for building, testing, and deploying high-performance machine learning systems on unbounded time-series data. Traditional data science workflows assume finite datasets and require substantial reimplementation when moving from batch prototypes to streaming production systems. This gap introduces causality violations, batch boundary artifacts, and poor reproducibility of real-time failures. DataFlow resolves these issues through a unified execution model based on directed acyclic graphs (DAGs) with point-in-time idempotency: outputs at any time t depend only on a fixed-length context window preceding t. This guarantee ensures that models developed in batch mode execute identically in streaming production without code changes. The framework enforces strict causality by automatically tracking knowledge time across all transformations, eliminating future-peeking bugs. DataFlow supports flexible tiling across temporal and feature dimensions, allowing the same model to operate at different frequencies and memory profiles via configuration alone. It integrates natively with the Python data science stack and provides fit/predict semantics for online learning, caching and incremental computation, and automatic parallelization through DAG-based scheduling. We demonstrate its effectiveness across domains including financial trading, IoT, fraud detection, and real-time analytics.

cs.LG

Causal Inference in Energy Demand Prediction

Energy demand prediction is critical for grid operators, industrial energy consumers, and service providers. Energy demand is influenced by multiple factors, including weather conditions (e.g. temperature, humidity, wind speed, solar radiation), and calendar information (e.g. hour of day and month of year), which further affect daily work and life schedules. These factors are causally interdependent, making the problem more complex than simple correlation-based learning techniques satisfactorily allow for. We propose a structural causal model that explains the causal relationship between these variables. A full analysis is performed to validate our causal beliefs, also revealing important insights consistent with prior studies. For example, our causal model reveals that energy demand responds to temperature fluctuations with season-dependent sensitivity. Additionally, we find that energy demand exhibits lower variance in winter due to the decoupling effect between temperature changes and daily activity patterns. We then build a Bayesian model, which takes advantage of the causal insights we learned as prior knowledge. The model is trained and tested on unseen data and yields state-of-the-art performance in the form of a 3.84 percent MAPE on the test set. The model also demonstrates strong robustness, as the cross-validation across two years of data yields an average MAPE of 3.88 percent.

cs.AI

Runnable Directories: The Solution to the Monorepo vs. Multi-repo Debate

Modern software systems increasingly strain traditional codebase organization strategies. Monorepos offer consistency but often suffer from scalability issues and tooling complexity, while multi-repos provide modularity at the cost of coordination and dependency management challenges. As an answer to this trade-off, we present the Causify Dev system, a hybrid approach that integrates key benefits of both. Its central concept is the runnable directory -- a self-contained, independently executable unit with its own development, testing, and deployment lifecycles. Backed by a unified thin environment, shared helper utilities, and containerized Docker-based workflows, runnable directories enable consistent setups, isolated dependencies, and efficient CI/CD processes. The Causify Dev approach provides a practical middle ground between monorepo and multi-repo strategies, improving reliability and maintainability for growing, complex codebases.

cs.SE

A Benchmark of Causal vs. Correlation AI for Predictive Maintenance

Predictive maintenance in manufacturing environments presents a challenging optimization problem characterized by extreme cost asymmetry, where missed failures incur costs roughly fifty times higher than false alarms. Predictive maintenance in manufacturing environments presents a challenging optimization problem characterized by extreme cost asymmetry, where missed failures incur costs roughly fifty times higher than false alarms. Conventional machine learning approaches typically optimize statistical accuracy metrics that do not reflect this operational reality and cannot reliably distinguish causal relationships from spurious correlations. This study benchmarks eight predictive models, ranging from baseline statistical approaches to Bayesian structural causal methods, on a dataset of 10,000 CNC machines with a 3.3 percent failure prevalence. While ensemble correlation-based models such as Random Forest (L4) achieve the highest raw cost savings (70.8 percent reduction), the Bayesian Structural Causal Model (L7) delivers competitive financial performance (66.4 percent cost reduction) with an inherent ability of failure attribution, which correlation-based models do not readily provide. The model achieves perfect attribution for HDF, PWF, and OSF failure types. These results suggest that causal methods, when combined with domain knowledge and Bayesian inference, offer a potentially favorable trade-off between predictive performance and operational interpretability in predictive maintenance applications.

cs.AI

The effect of latency on optimal order execution policy

Market participants regularly send bid and ask quotes to exchange-operated limit order books. This creates an optimization challenge where their potential profit is determined by their quoted price and how often their orders are successfully executed. The expected profit from successful execution at a favorable limit price needs to be balanced against two key risks: (1) the possibility that orders will remain unfilled, which hinders the trading agenda and leads to greater price uncertainty, and (2) the danger that limit orders will be executed as market orders, particularly in the presence of order submission latency, which in turn results in higher transaction costs. In this paper, we consider a stochastic optimal control problem where a risk-averse trader attempts to maximize profit while balancing risk. The market is modeled using Brownian motion to represent the price uncertainty. We analyze the relationship between fill probability, limit price, and order submission latency. We derive closed-form approximations of these quantities that perform well in the practical regime of interest. Then, we utilize a mean-variance method where our total reward function features a risk-tolerance parameter to quantify the combined risk and profit.

q-fin.MF

On the Effect of Alpha Decay and Transaction Costs on the Multi-period Optimal Trading Strategy

We consider the multi-period portfolio optimization problem with a single asset that can be held long or short. Due to the presence of transaction costs, maximizing the immediate reward at each period may prove detrimental, as frequent trading results in consistent negative cash outflows. To simulate alpha decay, we consider a case where not only the present value of a signal, but also past values, have predictive power. We formulate the problem as an infinite horizon Markov Decision Process and seek to characterize the optimal policy that realizes the maximum average expected reward. We propose a variant of the standard value iteration algorithm for computing the optimal policy. Establishing convergence in our setting is nontrivial, and we provide a rigorous proof. Addtionally, we compute a first-order approximation and asymptotics of the optimal policy with small transaction costs.

math.OC

The FlEye camera: Sampling the joint distribution of natural scenes and motion

To make efficient use of limited physical resources, the brain must match its coding and computational strategies to the statistical structure of input signals. An attractive testing ground for these principles is the problem of motion estimation in the fly visual system: we understand the optics of the compound eye, have a quantitative description of input signals and noise from the retina, and can record from output neurons that encode estimates of different velocity components. Furthermore, recent work provides a nearly complete wiring diagram of the intervening circuitry. What is missing is a characterization of the visual signals and motions that flies encounter in a natural context. We attack this directly with the development of a specialized camera that matches the high temporal resolution, optical properties, and spectral sensitivity of the fly's eye; inertial motion sensors provide ground truth about rotations and translations through the world. We describe the design, construction, and performance characteristics of this FlEye camera. To illustrate the opportunities created by this instrument we use data on movies and motion to construct optimal local motion estimators that can be compared with the responses of the fly's motion sensitive neurons.

q-bio.NC

Optimal two-parameter portfolio management strategy with transaction costs

We consider a simplified model for optimizing a single-asset portfolio in the presence of transaction costs given a signal with a certain autocorrelation and cross-correlation structure. In our setup, the portfolio manager is given two one-parameter controls to influence the construction of the portfolio. The first is a linear filtering parameter that may increase or decrease the level of autocorrelation in the signal. The second is a numerical threshold that determines a symmetric "no-trade" zone. Portfolio positions are constrained to a single unit long or a single unit short. These constraints allow us to focus on the interplay between the signal filtering mechanism and the hysteresis introduced by the "no-trade" zone. We then formulate an optimization problem where we aim to minimize the frequency of trades subject to a fixed return level of the portfolio. We show that maintaining a no-trade zone while removing autocorrelation entirely from the signal yields a locally optimal solution. For any given "no-trade" zone threshold, this locally optimal solution also achieves the maximum attainable return level, and we derive a quantitative lower bound for the amount of improvement in terms of the given threshold and the amount of autocorrelation removed.

math.OC

Towards a Systematic Approach for Smart Grid Hazard Analysis and Experiment Specification

The transition to the smart grid introduces complexity to the design and operation of electric power systems. This complexity has the potential to result in safety-related losses that are caused, for example, by unforeseen interactions between systems and cyber-attacks. Consequently, it is important to identify potential losses and their root causes, ideally during system design. This is non-trivial and requires a systematic approach. Furthermore, due to complexity, it may not possible to reason about the circumstances that could lead to a loss; in this case, experiments are required. In this work, we present how two complementary deductive approaches can be usefully integrated to address these concerns: Systems Theoretic Process Analysis (STPA) is a systems approach to identifying safety-related hazard scenarios; and the ERIGrid Holistic Test Description (HTD) provides a structured approach to refine and document experiments. The intention of combining these approaches is to enable a systematic approach to hazard analysis whose findings can be experimentally tested. We demonstrate the use of this approach with a reactive power voltage control case study for a low voltage distribution network.

cs.SE

Universality for two-dimensional critical cellular automata

We study the class of monotone, two-state, deterministic cellular automata, in which sites are activated (or 'infected') by certain configurations of nearby infected sites. These models have close connections to statistical physics, and several specific examples have been extensively studied in recent years by both mathematicians and physicists. This general setting was first studied only recently, however, by Bollobás, Smith and Uzzell, who showed that the family of all such 'bootstrap percolation' models on $\mathbb{Z}^2$ can be naturally partitioned into three classes, which they termed subcritical, critical and supercritical. In this paper we determine the order of the threshold for percolation (complete occupation) for every critical bootstrap percolation model in two dimensions. This 'universality' theorem includes as special cases results of Aizenman and Lebowitz, Gravner and Griffeath, Mountford, and van Enter and Hulshof, significantly strengthens bounds of Bollobás, Smith and Uzzell, and complements recent work of Balister, Bollobás, Przykucki and Smith on subcritical models.

math.PR

Subcritical monotone cellular automata

We study monotone cellular automata (also known as $\mathcal{U}$-bootstrap percolation) in $\mathbb{Z}^d$ with random initial configurations. Confirming a conjecture of Balister, Bollobás, Przykucki and Smith, who proved the corresponding result in two dimensions, we show that the critical probability is non-zero for all subcritical models.

math.PR

Universality for monotone cellular automata

In this paper we study monotone cellular automata in $d$ dimensions. We develop a general method for bounding the growth of the infected set when the initial configuration is chosen randomly, and then use this method to prove a lower bound on the critical probability for percolation that is sharp up to a constant factor in the exponent for every 'critical' model. This is one of three papers that together confirm the Universality Conjecture of Bollob\'as, Duminil-Copin, Morris and Smith.

math.PR

The critical length for growing a droplet

In many interacting particle systems, relaxation to equilibrium is thought to occur via the growth of 'droplets', and it is a question of fundamental importance to determine the critical length at which such droplets appear. In this paper we construct a mechanism for the growth of droplets in an arbitrary finite-range monotone cellular automaton on a $d$-dimensional lattice. Our main application is an upper bound on the critical probability for percolation that is sharp up to a constant factor in the exponent. Our method also provides several crucial tools that we expect to have applications to other interacting particle systems, such as kinetically constrained spin models on $\mathbb{Z}^d$. This is one of three papers that together confirm the Universality Conjecture of Bollob\'as, Duminil-Copin, Morris and Smith.

math.PR

On the Estimation of Entropy in the FastICA Algorithm

The fastICA method is a popular dimension reduction technique used to reveal patterns in data. Here we show both theoretically and in practice that the approximations used in fastICA can result in patterns not being successfully recognised. We demonstrate this problem using a two-dimensional example where a clear structure is immediately visible to the naked eye, but where the projection chosen by fastICA fails to reveal this structure. This implies that care is needed when applying fastICA. We discuss how the problem arises and how it is intrinsically connected to the approximations that form the basis of the computational efficiency of fastICA.

stat.ML

SN 2014ab: An Aspherical Type IIn Supernova with Low Polarization

We present photometry, spectra, and spectropolarimetry of supernova (SN) 2014ab, obtained through $\sim 200$ days after peak brightness. SN 2014ab was a luminous Type IIn SN ($M_V < -19.14$ mag) discovered after peak brightness near the nucleus of its host galaxy, VV 306c. Prediscovery upper limits constrain the time of explosion to within 200 days prior to discovery. While SN 2014ab declined by $\sim 1$ mag over the course of our observations, the observed spectrum remained remarkably unchanged. Spectra exhibit an asymmetric emission-line profile with a consistently stronger blueshifted component, suggesting the presence of dust or a lack of symmetry between the far side and near side of the SN. The Pa$β$ emission line shows a profile very similar to that of H$α$, implying that this stronger blueshifted component is caused either through obscuration by large dust grains, occultation by optically thick material, or a lack of symmetry between the far side and near side of the interaction region. Despite these asymmetric line profiles, our spectropolarimetric data show that SN 2014ab has little detected polarization after accounting for the interstellar polarization. This suggests that we are seeing emission from a photosphere that has only small deviation from circular symmetry face-on. We are likely seeing a SN IIn with nearly circular symmetry in the plane normal to our line of sight, but with either large-grain dust or significant asymmetry in the density of circumstellar material or SN ejecta along our line of sight. We suggest that SN 2014ab and SN 2010jl (as well as other SNe IIn) may be similar events viewed from different directions.

astro-ph.SR

What Makes Ly$α$ Nebulae Glow? Mapping the Polarization of LABd05

"Ly$α$ nebulae" are giant ($\sim$100 kpc), glowing gas clouds in the distant universe. The origin of their extended Ly$α$ emission remains a mystery. Some models posit that Ly$α$ emission is produced when the cloud is photoionized by UV emission from embedded or nearby sources, while others suggest that the Ly$α$ photons originate from an embedded galaxy or AGN and are then resonantly scattered by the cloud. At least in the latter scenario, the observed Ly$α$ emission will be polarized. To test these possibilities, we are conducting imaging polarimetric observations of seven Ly$α$ nebulae. Here we present our results for LABd05, a cloud at $z$ = 2.656 with an obscured, embedded AGN to the northeast of the peak of Ly$α$ emission. We detect significant polarization. The highest polarization fractions $P$ are $\sim$10-20% at $\sim$20-40 kpc southeast of the Ly$α$ peak, away from the AGN. The lowest $P$, including upper-limits, are $\sim$5% and lie between the Ly$α$ peak and AGN. In other words, the polarization map is lopsided, with $P$ increasing from the Ly$α$ peak to the southeast. The measured polarization angles $θ$ are oriented northeast, roughly perpendicular to the $P$ gradient. This unique polarization pattern suggests that 1) the spatially-offset AGN is photoionizing nearby gas and 2) escaping Ly$α$ photons are scattered by the nebula at larger radii and into our sightline, producing tangentially-oriented, radially-increasing polarization away from the photoionized region. Finally we conclude that the interplay between the gas density and ionization profiles produces the observed central peak in the Ly$α$ emission. This also implies that the structure of LABd05 is more complex than assumed by current theoretical spherical or cylindrical models.

astro-ph.GA