SearcharxivSearch

arXiv subjects

Yingjie Zhang

Publications and source records attributed to Yingjie Zhang.

At least 19 recordsLinked to original sources

The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions

Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action resembles an unauthorized one, and the safe-versus-unsafe label is a choice of authorization policy, not fixed by the action alone. We argue it therefore requires a benchmark that does not yet exist, one that maps the decision boundary of an ideal guardrail. Harvesting such a benchmark from real data is impractical: boundary cases are hard to collect, their labels hard to verify. The gap is real, so we construct Cautious Bench, the first benchmark to make over-safety the construct for agent guardrails; it codesigns each sample and its label with a stated authorization policy. A build-time gate re-derives every example to certify it, so each label is a mechanical consequence of the policy rather than an annotator's per-sample verdict, a reference against which researchers can measure real guardrails. The benchmark renders 756 Decidable benign/twin pairs, each under three object-name types (2,268 measured pairs), and 40 Undecidable pairs reported separately. Measuring six guardrails from five designs, we find a name-superstition effect: each over-refuses an authorized action more often under a scary-looking object name than a benign one. Since only the object name varies in the aforementioned contrast experiments, the deviation is the name's doing: the guardrails read the surface label, not the authorization context.

cs.CR

Variable-Path-Length FTIR of E. coli in Aqueous Media

Transmission infrared spectroscopy has been widely used for chemical analysis of biological samples in aqueous environments. However, its scope of applications has been limited by the path length, which is either too large or fixed, posing challenges for analyzing highly absorbing or heterogeneous samples. In this work, a mid-infrared (mid-IR) optical fiber probe was used for Fourier transform infrared (FTIR) micro-spectroscopy of aqueous Escherichia coli (E. coli) samples, providing continuous tuning of optical path length and sampling of near-surface and bulk regions. The mid-IR absorbance of the protein signal at 1548 cm-1 increased linearly with path length, consistent with the Beer-Lambert law. Path-length dependent spectra were used to calculate the spatially heterogeneous absorption coefficient of E. coli suspensions in aqueous media. The results demonstrate the ability of our fiber-based technique to resolve signals originating from different depths into the biological solution.

physics.ins-det

Oxygen Reduction Reaction on Platinum Nanocatalysts Produces Long-Lived, Hysteretic Oxygenated Adsorbates

Aqueous electrocatalysis generates oxygenated intermediates at catalyst surfaces. While intermediate species on single-crystal catalysts have been observed, the nature and evolution of surface oxygenated species on industrially relevant nanoparticle (NP) catalysts remain largely unknown. Here, using in situ Raman spectroscopy, we tracked the formation and potential-dependent evolution of oxygenated adsorbates in alkaline media on NP catalysts with an active platinum (Pt) surface. By comparing spectroscopic features in Ar- vs O2-saturated electrolytes, we determined three key intermediates produced by the oxygen reduction reaction (ORR): adsorbed OOH, OH, and O2. In contrast to the conventional wisdom that intermediates exist only during catalytic reactions, we found these oxygenated adsorbates to be highly long-lived and hysteretic, and to persist even after the termination of ORR. This adsorbate-retention effect exhibits a modest dependence on the surface oxidation state and the electrolyte cations (K+ vs Li+), and is likely facilitated by the heterogeneous nature of the catalyst surface. The results highlight the complexity of surface adsorption structures on realistic catalysts, which often extends beyond that captured by measurements or simulations on model single-crystal surfaces.

physics.chem-ph

What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection

Modern agentic systems fundamentally reshape the security boundary of LLMs by introducing persistent system state including memories, filesystems, tools, and other long-lived contextual artifacts that survives across sessions. As external information crosses this boundary and becomes part of persistent agent state, malicious instructions are no longer confined to a single interaction, but can silently persist and influence future executions long after the original attacker interaction has ended. We introduce Cross-Session Stored Prompt Injection, a new threat vector inspired by stored cross-site scripting that redefines prompt injection for agentic systems by extending its threat model across both time, where attacks persist and activate across sessions, and space, where adversarial instructions propagate beyond the immediate prompt into persistent system state. To systematically characterize this emerging threat, we formalize the lifecycle of cross-session stored prompt injection, develop a taxonomy of persistence channels and incorporation mechanisms, and build a sandbox toolkit for evaluation. Our findings suggest that the fundamental challenge of agent security is not merely filtering untrusted inputs, but governing how external information acquires authority as it crosses persistent system boundaries. We hope this work motivates a broader shift from interaction-centric security toward state-centric security, making the secure management of persistent agent state a first-class security principle for the agentic era.

cs.CR

The XMM-Newton Line Emission Analysis Program (X-LEAP) III: Earth's Magnetospheric X-ray Emission Revealed by 22-Year XMM-Newton Observations

The magnetosphere, protecting the Earth from intense solar activity, is also shaped by the solar wind, while its structure is still uncertain in observation. In this study, we map the X-ray emission in the magnetosphere, which is induced by the charge exchange between the highly-ionized solar wind and the neutral gas around the Earth, known as the magnetospheric solar wind charge exchange (SWCX). In particular, we extract the magnetospheric SWCX in the O VII line emission data adopted from the XMM Line Emission Analysis Program (X-LEAP). The observed magnetospheric SWCX shows an enhanced emission of approximately $I_{\rm OVII}^{\rm mag}\approx 2$ photons $\rm cm^{-2}~ s^{-1}~sr^{-1}$ toward the Sun, showing a consistent shape predicted by numerical simulations. Furthermore, this magnetospheric SWCX exhibits a dependence on the XMM pointing direction, which traces the path length of SWCX emission in the Earth's magnetosphere. Building on this directional dependence, we model the 3D magnetosheath structure using soft X-ray observations for the first time, constraining the averaged boundary geometry and SWCX emissivity distribution over 22 years. Finally, utilizing the XMM data, we derive an empirical O VII emission efficiency of $α_{\rm OVII}=(2.1\pm0.4) \times 10^{-16}\ {\rm eV\,cm^{2}}$.

astro-ph.HE

CRT*: Conditional Randomization Testing with Heterogeneous External and Unlabeled Data

The conditional randomization test (CRT) provides a principled approach to conditional independence (CI) testing, guaranteeing exact type-I error control when the true conditional distribution is known. In practice, however, this distribution must be estimated, and estimation errors can inflate type-I errors, while high dimensionality and limited sample sizes can reduce power. Although external and unlabeled data offer the potential to improve CI testing, naive integration that ignores distributional heterogeneity can compromise type-I error control and fail to enhance power. We propose \textbf{CRT*}, a novel framework that robustly integrates external and unlabeled datasets to enhance CI testing in heterogeneous scenarios. CRT* employs smooth residual-bootstrap (SRB) with transfer learning for conditional distribution estimation, combined with adaptive data fusion via an optimal convex combination of test statistics. We theoretically establish that the SRB-based estimator converges to the true conditional distribution in expected total variation distance. Furthermore, even in high-dimensional regimes, CRT* maintains valid type-I error control and achieves strictly higher power than standard CRT without external data. Simulations and RNA-seq breast cancer data analyses demonstrate that CRT* substantially improves power while maintaining type-I error control in heterogeneous settings.

stat.ME

A Diagnostic Framework for AI Agent Behavior

AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.

cs.AI

NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models

Out-of-Distribution (OOD) detection is essential for ensuring the robustness and reliability of object detection systems deployed in safety-critical applications. While prior research has mainly focused on uni-modal detectors or vision-language model (VLM) based classifiers, the potential of VLM-based object detectors in OOD scenarios remains underexplored. In this work, we take the first step toward building OOD object detection methods upon VLMs. We identify two challenges specific to VLM detectors: (i) their text-guided attention enhances foreground with ID labels but treats background uniformly, leaving potential OOD regions unexploited for separating in-distribution (ID) from OOD instances; and (ii) their sigmoid-based multi-label outputs are incompatible with softmax-based OOD scores, calling for scoring functions consistent with VLM probabilistic outputs. Hence, we introduce Negative Label Guided Attention and Scoring (NegAS). To address (i), we propose a negative label guided attention module (NegA), where LLM-generated, visually-similar but semantically-different negative labels are used to guide attention toward potential OOD background regions. To address (ii), we introduce a novel sigmoid-based OOD scoring function (NegS) that leverages both ID and negative labels, producing strong responses for ID instances and suppressed responses for OOD ones. Extensive experiments demonstrate that our approach improves OOD detection performance by a large margin while maintaining ID accuracy, e.g., reducing the FPR95 by 11.4% on the COCO dataset and 25.5% on the OpenImages dataset compared to the baseline model. While initially designed for dense VLM detectors like YOLO-World, we successfully adapt NegAS to Grounding DINO, a query-based VLM transformer and achieve significant improvements, demonstrating the generalizability of our framework.

cs.CV

Agents for Experiments, Experiments for Agents: A Design Grammar for AI-Enabled Experimental Science

AI systems are becoming active participants in organizational and knowledge work. They increasingly interact with humans, coordinate workflows, and operate in multi-agent arrangements. Understanding their effects therefore requires more than measuring output accuracy; it requires evidence about mechanisms, delegation, feedback, and control. Experiments remain central to this task, but they also face a recursive challenge: we need experiments for agents to study these arrangements, and we may need agents for experiments to help search the expanding space of possible designs. Yet experimental conditions for human-AI and agentic workflows are still largely specified in prose, making them difficult to compare, reuse, or audit. We frame this as a problem of workflow representation, traceability, and governance in AI-enabled knowledge production. We introduce SEED (Structural Encoding for Experimental Discovery), a framework that represents experimental conditions as typed actor-flow graphs. SEED supports three design functions: describing conditions as interaction structures, evaluating structural novelty relative to encoded prior designs, and generating candidate designs under feasibility and governance constraints. We report a lightweight empirical feasibility test that compares graph-blind and SEEDguided generation in a medical-triage design task. In this diagnostic contrast, SEED-guided candidate designs show clearer actor-flow changes, assumptions, and governance checks, supporting the feasibility of the grammar as a design aid. The commentary closes by identifying governance tensions around novelty, replication, validity, diversity of inquiry, and accountability.

cs.AI

Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game

Retrieval-Augmented Generation (RAG) systems augment large language models with external knowledge, yet introduce a critical security vulnerability: RAG Knowledge Base Leakage, wherein adversarial prompts can induce the model to divulge retrieved proprietary content. Recent studies reveal that such leakage can be executed through adaptive and iterative attack strategies (named RAG extraction attack), while effective countermeasures remain notably lacking. To bridge this gap, we propose CanaryRAG, a runtime defense mechanism inspired by stack canaries in software security. CanaryRAG embeds carefully designed canary tokens into retrieved chunks and reformulates RAG extraction defense as a dual-path runtime integrity game. Leakage is detected in real time whenever either the target or oracle path violates its expected canary behavior, including under adaptive suppression and obfuscation. Extensive evaluations against existing attacks demonstrate that CanaryRAG provides robust defense, achieving substantially lower chunk recovery rates than state-of-the-art baselines while imposing negligible impact on task performance and inference latency. Moreover, as a plug-and-play solution, CanaryRAG can be seamlessly integrated into arbitrary RAG pipelines without requiring retraining or structural modifications, offering a practical and scalable safeguard for proprietary data.

cs.CR

Liquid structure adjacent to solid surfaces follows the superposition principle

Liquid structure at solid-liquid interfaces is critical for many natural and engineered processes ranging from biological signal transduction to electrochemical energy conversion. Advanced experimental and computational methods have provided insights into the structure of liquids adjacent to planar substrates at the nanoscale. However, realistic solid-liquid interfaces are inevitably inhomogeneous across multiple length scales, presenting a complexity that surpasses the capabilities of existing approaches. Here we bridge the complexity gap by discovering and utilizing a hitherto hidden principle of interfacial liquid--superposition. Experimentally, we use 3D atomic force microscopy (3D-AFM) to image the interfacial structure of a wide range of organic and aqueous solvents and electrolytes, uncovering universal liquid density oscillations and emergent liquid layer reconfigurations at heterogeneous substrate sites. We further develop an analytical model, coined solid-liquid superposition (SLS), which solves the interfacial liquid density distribution based on a key descriptor: the effective total correlation function (ETCF) between a liquid molecule and nearby solid atoms. SLS not only explains all the experimentally observed interfacial liquid distribution profiles from the angstrom to near-micron scale, but also predicts more precise atomic-scale interference patterns which are further corroborated by molecular dynamics (MD) simulations. This study unveils a key structural descriptor of interfacial liquids, and establishes a theoretical framework for rapidly and accurately predicting liquid structures adjacent to solid surfaces with arbitrary morphology and size scale.

physics.chem-ph

Visioning Human-Agentic AI Teaming: Continuity, Tension, and Future Research

Artificial intelligence is undergoing a structural transformation marked by the rise of agentic systems capable of open-ended action trajectories, generative representations and outputs, and evolving objectives. These properties introduce structural uncertainty into human-AI teaming (HAT), including uncertainty about behavior trajectories, epistemic grounding, and the stability of governing logics over time. Under such conditions, alignment cannot be secured through agreement on bounded outputs; it must be continuously sustained as plans unfold and priorities shift. We advance Team Situation Awareness (Team SA) theory, grounded in shared perception, comprehension, and projection, as an integrative anchor for this transition. While Team SA remains analytically foundational, its stabilizing logic presumes that shared awareness, once achieved, will support coordinated action through iterative updating. Agentic AI challenges this presumption. Our argument unfolds in two stages: first, we extend Team SA to reconceptualize both human and AI awareness under open-ended agency, including the sensemaking of projection congruence across heterogeneous systems. Second, we interrogate whether the dynamic processes traditionally assumed to stabilize teaming in relational interaction, cognitive learning, and coordination and control continue to function under adaptive autonomy. By distinguishing continuity from tension, we clarify where foundational insights hold and where structural uncertainty introduces strain, and articulate a forward-looking research agenda for HAT. The central challenge of HAT is not whether humans and AI can agree in the moment, but whether they can remain aligned as futures are continuously generated, revised, enacted, and governed over time.

cs.AI

Conditionally Whitened Generative Models for Probabilistic Time Series Forecasting

Probabilistic forecasting of multivariate time series is challenging due to non-stationarity, inter-variable dependencies, and distribution shifts. While recent diffusion and flow matching models have shown promise, they often ignore informative priors such as conditional means and covariances. In this work, we propose Conditionally Whitened Generative Models (CW-Gen), a framework that incorporates prior information through conditional whitening. Theoretically, we establish sufficient conditions under which replacing the traditional terminal distribution of diffusion models, namely the standard multivariate normal, with a multivariate normal distribution parameterized by estimators of the conditional mean and covariance improves sample quality. Guided by this analysis, we design a novel Joint Mean-Covariance Estimator (JMCE) that simultaneously learns the conditional mean and sliding-window covariance. Building on JMCE, we introduce Conditionally Whitened Diffusion Models (CW-Diff) and extend them to Conditionally Whitened Flow Matching (CW-Flow). Experiments on five real-world datasets with six state-of-the-art generative models demonstrate that CW-Gen consistently enhances predictive performance, capturing non-stationary dynamics and inter-variable correlations more effectively than prior-free approaches. Empirical results further demonstrate that CW-Gen can effectively mitigate the effects of distribution shift.

stat.ML

A Synthetic Data-Driven Radiology Foundation Model for Pan-tumor Clinical Diagnosis

AI-assisted imaging made substantial advances in tumor diagnosis and management. However, a major barrier to developing robust oncology foundation models is the scarcity of large-scale, high-quality annotated datasets, which are limited by privacy restrictions and the high cost of manual labeling. To address this gap, we present PASTA, a pan-tumor radiology foundation model built on PASTA-Gen, a synthetic data framework that generated 30,000 3D CT scans with pixel-level lesion masks and structured reports of tumors across ten organ systems. Leveraging this resource, PASTA achieves state-of-the-art performance on 45 of 46 oncology tasks, including non-contrast CT tumor screening, lesion segmentation, structured reporting, tumor staging, survival prediction, and MRI-modality transfer. To assess clinical applicability, we developed PASTA-AID, a clinical decision support system, and ran a retrospective simulated clinical trial across two scenarios. For pan-tumor screening on plain CT with fixed reading time, PASTA-AID increased radiologists' throughput by 11.1-25.1% and improved sensitivity by 17.0-31.4% and precision by 10.5-24.9%; additionally, in a diagnosis-aid workflow, it reduced segmentation time by up to 78.2% and reporting time by up to 36.5%. Beyond gains in accuracy and efficiency, PASTA-AID narrowed the expertise gap, enabling less-experienced radiologists to approach expert-level performance. Together, this work establishes an end-to-end, synthetic data-driven pipeline spanning data generation, model development, and clinical validation, thereby demonstrating substantial potential for pan-tumor research and clinical translation.

eess.IV

PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognition (OCR) in the low-resource Pashto language. Natural Language Processing (NLP) in Pashto faces several challenges due to the cursive nature of its script and a scarcity of structured datasets. To address this, we developed a synthetic Pashto OCR dataset, PsOCR, consisting of one million images annotated with bounding boxes at word, line, and document levels, suitable for training and evaluating models based on different architectures, including Convolutional Neural Networks (CNNs) and Transformers. PsOCR covers variations across 1,000 unique font families, colors, image sizes, and layouts. A benchmark subset of 10K images was selected to evaluate the performance of several LMMs, including seven open-source models: DeepSeek's Janus, InternVL, MiniCPM, Florence, and Qwen (3B and 7B), and four closed-source models: GPT-4o, Gemini, Claude, and Grok. Experimental results demonstrate that Gemini achieves the best performance among all models, whereas among open-source models, Qwen-7B stands out. This work provides an insightful assessment of the capabilities and limitations of current LMMs for OCR tasks in Pashto and establishes a foundation for further research not only in Pashto OCR but also for other similar scripts such as Arabic, Persian, and Urdu. PsOCR is available at https://github.com/zirak-ai/PashtoOCR.

cs.CV

Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks

LLMs remain vulnerable to jailbreak attacks that exploit adversarial prompts to circumvent safety measures. Current safety fine-tuning approaches face two critical limitations. First, they often fail to strike a balance between security and utility, where stronger safety measures tend to over-reject harmless user requests. Second, they frequently miss malicious intent concealed within seemingly benign tasks, leaving models exposed to exploitation. Our work identifies a fundamental cause of these issues: during response generation, an LLM's capacity to differentiate harmful from safe outputs deteriorates. Experimental evidence confirms this, revealing that the separability between hidden states for safe and harmful responses diminishes as generation progresses. This weakening discrimination forces models to make compliance judgments earlier in the generation process, restricting their ability to recognize developing harmful intent and contributing to both aforementioned failures. To mitigate this vulnerability, we introduce DEEPALIGN - an inherent defense framework that enhances the safety of LLMs. By applying contrastive hidden-state steering at the midpoint of response generation, DEEPALIGN amplifies the separation between harmful and benign hidden states, enabling continuous intrinsic toxicity detection and intervention throughout the generation process. Across diverse LLMs spanning varying architectures and scales, it reduced attack success rates of nine distinct jailbreak attacks to near-zero or minimal. Crucially, it preserved model capability while reducing over-refusal. Models equipped with DEEPALIGN exhibited up to 3.5% lower error rates in rejecting challenging benign queries and maintained standard task performance with less than 1% decline. This marks a substantial advance in the safety-utility Pareto frontier.

cs.CR

A benchmark multimodal oro-dental dataset for large vision-language models

The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset, comprising 8775 dental checkups from 4800 patients collected over eight years (2018-2025), with patients ranging from 10 to 90 years of age. The dataset includes 50000 intraoral images, 8056 radiographs, and detailed textual records, including diagnoses, treatment plans, and follow-up notes. The data were collected under standard ethical guidelines and annotated for benchmarking. To demonstrate its utility, we fine-tuned state-of-the-art large vision-language models, Qwen-VL 3B and 7B, and evaluated them on two tasks: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. We compared the fine-tuned models with their base counterparts and GPT-4o. The fine-tuned models achieved substantial gains over these baselines, validating the dataset and underscoring its effectiveness in advancing AI-driven oro-dental healthcare solutions. The dataset is publicly available, providing an essential resource for future research in AI dentistry.

cs.CV

Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction

Jailbreak attacks pose persistent threats to large language models (LLMs). Current safety alignment methods have attempted to address these issues, but they experience two significant limitations: insufficient safety alignment depth and unrobust internal defense mechanisms. These limitations make them vulnerable to adversarial attacks such as prefilling and refusal direction manipulation. We introduce DeepRefusal, a robust safety alignment framework that overcomes these issues. DeepRefusal forces the model to dynamically rebuild its refusal mechanisms from jailbreak states. This is achieved by probabilistically ablating the refusal direction across layers and token depths during fine-tuning. Our method not only defends against prefilling and refusal direction attacks but also demonstrates strong resilience against other unseen jailbreak strategies. Extensive evaluations on four open-source LLM families and six representative attacks show that DeepRefusal reduces attack success rates by approximately 95%, while maintaining model capabilities with minimal performance degradation.

cs.CR