arXiv · 2603.10807
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
Abstract
Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Services, and Insurance (BFSI) deployments exposed to failures elicited through legally or professionally plausible framing. We introduce RAHS (Risk-Adjusted Harm Score), a risk-sensitive metric jointly capturing disclosure severity, disclaimer mitigation, and inter-judge agreement, and FinRedTeamBench, a 989-prompt benchmark spanning seven BFSI risk areas and 34 sub-categories mapped to regulatory frameworks. Evaluation uses an ensemble of three heterogeneous LLM judges, validated against human experts, and an adaptive multi-turn red-teaming pipeline. On nine open-weight models, RAHS preserves separation under near-ceiling ASR, ranking is stable under hyperparameter sweeps, and multi-turn pressure drives not only more jailbreaks but more operationally severe disclosures, exposing failure modes that single-turn, domain-agnostic evaluations cannot reveal.
Explore related subjects
Keep this discovery
Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali. 2026-08-31. Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services. https://arxiv.org/abs/2603.10807
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.