Loyal Agents: Training LLM Agents to Protect Principal Interests Under Strategic Information Asymmetry
As LLMs increasingly act as delegated agents, they are expected to protect principals' interests when interacting with external parties. Standard alignment objectives, such as helpfulness, harmlessness, and honesty, do not specify how agents should protect principals' strategic interests under delegation. We formalize Agent Loyalty as an information-control property requiring agents to prevent Exploitable Information Leakage (EIL) and resist Manipulative Information Uptake (MIU). We introduce LoyalAgent-Bench, comprising 10,298 samples across 42 subscenarios and six domains, and an online GRPO framework that trains against a LLM opponent to generate mechanism-specific reward signals. Experiments show that loyalty is not guaranteed by general capability or existing alignment, with measurable EIL and MIU gaps under zero-shot evaluation, while our trained 8B models reduce per-response leakage in single-turn exchanges by 31-44pp and improve task utility, evidence faithfulness, and decision accuracy by up to 11pp, 49pp, and 29pp, respectively. For the trained Qwen3-4B model, no degradation is observed on out-of-distribution benchmarks in math and narrative reasoning.