SearcharxivSearch

arXiv subjects

Eunna Lee

Publications and source records attributed to Eunna Lee.

4 recordsLinked to original sources

Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs

We ask whether a model protects a user in the same way when that user speaks rather than types. Using a single distress vignette---a physical injury of unstated severity following an interpersonal conflict---we present four frontier models with matched inputs across voice, text, and raw API deployment conditions (n=30 per cell) and code each response along five binary protective indicators, including whether the model issues an explicit medical-care directive. Voice-interface responses are markedly shorter than text-interface responses for three of the four models, and protective behavior contracts alongside that compression: medical directives are at ceiling under both the API and text conditions but decline under voice for every model tested. The contraction is not reducible to length. One model produces voice and text responses of comparable length yet still drops medical directives, and another falls below ceiling between its API and voice conditions, whose responses are of nearly identical length. Under raw API access the pattern is categorical rather than partial: no model asks after the user's safety even once. These results show that protective intervention is sensitive to the surface through which a request arrives, that this sensitivity is detectable using a simple protective coding scheme, and that it is not explained by turn length alone.

cs.CR

The Authority Expectancy Effect in Multi-Party Conflict

In multi-party competitive settings, across experiments on resource allocation, fault attribution, and dispute mediation, we test whether social authority (SA) cues such as occupational status and institutional documentation behave as if scalar weights were added to one side of a judgment. We reject this scalar-weight null on both of its predictions: cue effects are not independent in allocation decisions, where the same pair of cues interacts in opposite directions across models, and evidentiary cues do not exert a fixed directional influence in multi-turn disputes, where identical documentation draws protection to one party or extends it to both depending on which party holds it. The cues also govern whether a judgment is issued at all: removing an occupational label while retaining documentary evidence drives four of six models into refusal, yet the same label sharply reduces refusal when documentation is present, so the gate is subject to the same interaction as the judgment. We formalize this pattern as the Authority Expectancy Effect (AEE), comprising evidential reinterpretation, whereby identical content yields different judgments depending on which party bears the SA signal, and \textit{direction sensitivity}, whereby the same evidentiary cue favours different parties, or one party versus both, depending on the holder's authority position.

cs.AI

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives -- each preserving one objective at the cost of the other. Administering a three-turn escalating vulnerability vignette to three commercial LLMs (900 sessions across material, relational, and somatic status-proxy variants) and coding responses with two binary indices (VCC/VCI), we characterize a previously undocumented failure mode we term adaptive capitulation: the model validates the social injustice underlying the user's distress before pivoting to detailed facilitation of the very acquisition it nominally discouraged. We show that the trilemma is structural rather than incidental, and propose Minimal Reattributive Sufficiency (MRS), an architecture-neutral design principle that embeds a single reattributive cue within an otherwise validating response, preserving a pathway toward autonomous reattribution without contesting the user's stated goal.

cs.CL

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken, or to be taking, a real-world protective action it cannot perform, such as contacting emergency services or administering care. We term this phenomenon Protective Capacity Hallucination (PCH): a self-referential misattribution in which a model, acting in a protective role, asserts physical or institutional agency exceeding its affordances as a language model. In a three-phase study spanning eight LLMs and 13,600 sessions, we find that PCH depends on both situational severity and interactional format. Across ordinary service domains, multi-party dialogic input drives PCH to near-ceiling levels in most models. In contrast, PCH remains at floor levels in all eight models when the same models are placed in intimate-partner conflict scenarios, despite the greater physical severity of those situations. We interpret PCH as the signature of a deployment-design gap between role assignment and capability-boundary specification: a by-product of partial alignment in which a universally trained pressure to help outruns a domain-selective specification of how to help. Because suppression tracks alignment coverage rather than severity, deployment-side specification of capability boundaries emerges as a general mitigation target.

cs.CR