SearcharxivSearch

arXiv subjects

Austin Woo

Publications and source records attributed to Austin Woo.

2 recordsLinked to original sources

The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies

The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model differences significant. DeepSeek and Qwen are the most likely to recommend escalatory action using nuclear weapons; GPT and ERNIE are the least likely. Llama exhibits a distinct bias for action, favoring force, intervention, and cooperation across domains. Inter-rater reliability metrics (Krippendorff's $\alpha$ and quadratically weighted Fleiss' $\kappa$) reveal Llama and ERNIE are the most consistent across runs, with either DeepSeek or GLM the least depending on the domain. We also present a deeper exploration of our scenario variants: (i)~country-level biases tend to exist and vary by model, with country covariates like adversary trade ties and escalation propensity producing weak correlations; (ii)~existential phrasing effects are significant and heterogeneous; (iii)~these country biases interact with phrasing. Overall, the distributions of responses related to the scenarios in our benchmark vary significantly by model, country, and phrasing.

cs.CY

Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models

As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a novel benchmark designed to evaluate the biases and preferences of seven prominent foundation models-Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, GPT-4o, Gemini 1.5 Pro-002, Mixtral 8x22B, Claude 3.5 Sonnet, and Qwen2 72B-in the context of international relations (IR). We designed a bias discovery study around core topics in IR using 400-expert crafted scenarios to analyze results from our selected models. These scenarios focused on four topical domains including: military escalation, military and humanitarian intervention, cooperative behavior in the international system, and alliance dynamics. Our analysis reveals noteworthy variation among model recommendations based on scenarios designed for the four tested domains. Particularly, Qwen2 72B, Gemini 1.5 Pro-002 and Llama 3.1 8B Instruct models offered significantly more escalatory recommendations than Claude 3.5 Sonnet and GPT-4o models. All models exhibit some degree of country-specific biases, often recommending less escalatory and interventionist actions for China and Russia compared to the United States and the United Kingdom. These findings highlight the necessity for controlled deployment of LLMs in high-stakes environments, emphasizing the need for domain-specific evaluations and model fine-tuning to align with institutional objectives.

cs.CY