TY - RPRT TI - Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model AU - Pankayaraj Pathmanathan AU - Furong Huang PY - 2026 UR - https://arxiv.org/abs/2604.09665 ID - 2604.09665 ER -