SearcharxivSearch

arXiv subjects

Martin Pollack

Publications and source records attributed to Martin Pollack.

4 recordsLinked to original sources

The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies

The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model differences significant. DeepSeek and Qwen are the most likely to recommend escalatory action using nuclear weapons; GPT and ERNIE are the least likely. Llama exhibits a distinct bias for action, favoring force, intervention, and cooperation across domains. Inter-rater reliability metrics (Krippendorff's $\alpha$ and quadratically weighted Fleiss' $\kappa$) reveal Llama and ERNIE are the most consistent across runs, with either DeepSeek or GLM the least depending on the domain. We also present a deeper exploration of our scenario variants: (i)~country-level biases tend to exist and vary by model, with country covariates like adversary trade ties and escalation propensity producing weak correlations; (ii)~existential phrasing effects are significant and heterogeneous; (iii)~these country biases interact with phrasing. Overall, the distributions of responses related to the scenarios in our benchmark vary significantly by model, country, and phrasing.

cs.CY

CausalSent: Interpretable Sentiment Classification with RieszNet

Despite the overwhelming performance improvements offered by recent natural language processing (NLP) models, the decisions made by these models are largely a black box. Towards closing this gap, the field of causal NLP combines causal inference literature with modern NLP models to elucidate causal effects of text features. We replicate and extend Bansal et al's work on regularizing text classifiers to adhere to estimated effects, focusing instead on model interpretability. Specifically, we focus on developing a two-headed RieszNet-based neural network architecture which achieves better treatment effect estimation accuracy. Our framework, CausalSent, accurately predicts treatment effects in semi-synthetic IMDB movie reviews, reducing MAE of effect estimates by 2-3x compared to Bansal et al's MAE on synthetic Civil Comments data. With an ensemble of validated models, we perform an observational case study on the causal effect of the word "love" in IMDB movie reviews, finding that the presence of the word "love" causes a +2.9% increase in the probability of a positive sentiment.

cs.CL

Unraveling the Dynamics of SPY Trading Volumes: A Comprehensive Analysis of Daily and Intraday Liquidity Trends

In this project, we investigate the accuracy of forecasting intraday and daily trading volume of the exchange-traded fund SPY. The ability to forecast volume over varying time intervals with high accuracy is a critical element to many trading strategies. After performing exploratory data analysis on intraday and daily SPY data we identify three methods for our analysis: ARIMA and ARIMAX models, with or without seasonality, as well as a Frequency Domain Process Representation. To evaluate predictive power of our models, we use mean squared error, mean absolute percentage error, and volume weighted average price (VWAP) tracking error. All models for both intraday and daily data output strong VWAP predictions in comparison to the VWAP estimates produced by naive baseline methodologies. In both cases volume is most accurately forecasted using ARIMA models with exogenous variables in the form of technical indicators, with intraday incorporating a seasonal component and daily not.

stat.AP

Evaluation of Quadrature-based Moment Methods in turbulent premixed combustion

Transported probability density function (PDF) methods are widely used to model turbulent flames characterized by strong turbulence-chemistry interactions. Numerical methods directly resolving the PDF are commonly used, such as the Lagrangian particle or the stochastic fields (SF) approach. However, especially for premixed combustion configurations, characterized by high reaction rates and thin reaction zones, a fine PDF resolution is required, both in physical and in composition space, leading to high numerical costs. An alternative approach to solve a PDF is the method of moments, which has shown to be numerically efficient in a wide range of applications. In this work, two Quadrature-based Moment closures are evaluated in the context of turbulent premixed combustion. The Quadrature-based Moment Methods (QMOM) and the recently developed Extended QMOM (EQMOM) are used in combination with a tabulated chemistry approach to approximate the composition PDF. Both closures are first applied to an established benchmark case for PDF methods, a plug-flow reactor with imperfect mixing, and compared to reference results obtained from Lagrangian particle and SF approaches. Second, a set of turbulent premixed methane-air flames are simulated, varying the Karlovitz number and the turbulent length scale. The turbulent flame speeds obtained are compared with SF reference solutions. Further, spatial resolution requirements for simulating these premixed flames using QMOM are investigated and compared with the requirements of SF. The results demonstrate that both QMOM and EQMOM approaches are well suited to reproduce the turbulent flame properties. Additionally, it is shown that moment methods require lower spatial resolution compared to SF method.

physics.flu-dyn