SearcharxivSearch

arXiv subjects

Dominik Schwarz

Publications and source records attributed to Dominik Schwarz.

7 recordsLinked to original sources

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model families, an activation sensor blocks 95.5-97.7 percent of judge-classified compliant attacks in a taxonomy-selected set. It also blocks 59.6-68.4 percent of XSTest prompts. A fully disjoint audit reconstructs near-ceiling source-contrast AUROC (0.996-0.999), but fixed transfer to matched pairs is weaker: 0.656-0.819 on the guard-selected Twin-n70 subset and 0.590-0.690 on the full Twin-n163 cohort. We test ten axes on the reference family and seven across all families with leakage, hold-out, and permutation controls. On Twin-n163, no axis evaluated without direct pair-boundary fitting reaches the specified numerical threshold. Requiring persistence on that full cohort was added at analysis time. A separately specified 24B/32B extension gives the same result. Pair-trained classifiers weaken under category and generation-batch hold-out and false-block 79.6-100 percent of XSTest at 95 percent in-corpus TPR. At the tested read points, these activation scores behave as broad-risk detectors rather than standalone context adjudicators.

cs.CR

Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures

As Large Language Models (LLMs) are increasingly integrated into automated, multi-stage pipelines, risk patterns that arise from unvalidated trust between processing stages become a practical concern. This paper presents a mechanism-centered taxonomy of 41 recurring risk patterns in commercial LLMs. The analysis shows that inputs are often interpreted non-neutrally and can trigger implementation-shaped responses or unintended state changes even without explicit commands. We argue that these behaviors constitute architectural failure modes and that string-level filtering alone is insufficient. To mitigate such cross-stage vulnerabilities, we recommend zero-trust architectural principles, including provenance enforcement, context sealing, and plan revalidation, and we introduce "Countermind" as a conceptual blueprint for implementing these defenses.

cs.CR

Countermind: A Multi-Layered Security Architecture for Large Language Models

The security of Large Language Model (LLM) applications is fundamentally challenged by "form-first" attacks like prompt injection and jailbreaking, where malicious instructions are embedded within user inputs. Conventional defenses, which rely on post hoc output filtering, are often brittle and fail to address the root cause: the model's inability to distinguish trusted instructions from untrusted data. This paper proposes Countermind, a multi-layered security architecture intended to shift defenses from a reactive, post hoc posture to a proactive, pre-inference, and intra-inference enforcement model. The architecture proposes a fortified perimeter designed to structurally validate and transform all inputs, and an internal governance mechanism intended to constrain the model's semantic processing pathways before an output is generated. The primary contributions of this work are conceptual designs for: (1) A Semantic Boundary Logic (SBL) with a mandatory, time-coupled Text Crypter intended to reduce the plaintext prompt injection attack surface, provided all ingestion paths are enforced. (2) A Parameter-Space Restriction (PSR) mechanism, leveraging principles from representation engineering, to dynamically control the LLM's access to internal semantic clusters, with the goal of mitigating semantic drift and dangerous emergent behaviors. (3) A Secure, Self-Regulating Core that uses an OODA loop and a learning security module to adapt its defenses based on an immutable audit log. (4) A Multimodal Input Sandbox and Context-Defense mechanisms to address threats from non-textual data and long-term semantic poisoning. This paper outlines an evaluation plan designed to quantify the proposed architecture's effectiveness in reducing the Attack Success Rate (ASR) for form-first attacks and to measure its potential latency overhead.

cs.CR

Giant radio galaxies in the LoTSS Boötes deep field

Giant radio galaxies (GRGs) are radio galaxies that have projected linear extents of more than 700 kpc or 1 Mpc, depending on definition. We have carried out a careful visual inspection in search of GRGs of the Bootes LOFAR Deep Field (BLDF) image at 150 MHz. We identified 74 GRGs with a projected size larger than 0.7 Mpc of which 38 are larger than 1 Mpc. The resulting GRG sky density is about 2.8 (1.43) GRGs per square degree for GRGs with linear size larger than 0.7 (1) Mpc. We studied their radio properties and the accretion state of the host galaxies using deep optical and infrared survey data and determined flux densities for these GRGs from available survey images at both 54 MHz and 1.4 GHz to obtain integrated radio spectral indices. We show the location of the GRGs in the P-D diagram. The accretion mode onto the central black holes of the GRG hosts is radiatively inefficient suggesting that the central engines are not undergoing massive accretion at the time of the emission. Interestingly, 14 out of 35 GRGs for which optical spectra are available show a moderate star formation rate. Based on the number density of optical galaxies taken from the DESI DR9 photometric redshift catalogue, we found no significant differences between the environments of GRGs and other radio galaxies, at least for redshift up to z = 0.7.

astro-ph.GA

Cosmology with SKA Radio Continuum Surveys

Radio continuum surveys have, in the past, been of restricted use in cosmology. Most studies have concentrated on cross-correlations with the cosmic microwave background to detect the integrated Sachs-Wolfe effect, due to the large sky areas that can be surveyed. As we move into the SKA era, radio continuum surveys will have sufficient source density and sky area to play a major role in cosmology on the largest scales. In this chapter we summarise the experiments that can be carried out with the SKA as it is built up through the coming decade. We show that the SKA can play a unique role in constraining the non-Gaussianity parameter to σ(f_NL) ~ 1, and provide a unique handle on the systematics that inhibit weak lensing surveys. The SKA will also provide the necessary data to test the isotropy of the Universe at redshifts of order unity and thus evaluate the robustness of the cosmological principle.Thus, SKA continuum surveys will turn radio observations into a central probe of cosmological research in the coming decades.

astro-ph.CO

The Oddly Quiet Universe: How the CMB challenges cosmology's standard model

We discuss selected large-scale anomalies in the maps of temperature anisotropies in the cosmic microwave background. Specfically, these include alignments of the largest modes of CMB anisotropy with one another and with the geometry and direction of motion of the Solar System, and the unexpected absence of two-point angular corellations especially outside the region of the sky most contaminated by the Galaxy. We discuss these findings in relation to expectations from standard inflationary cosmology. This paper is adapted from a talk given by one of us (GDS) at the SEENET-2011 meeting in August 2011 on the Serbian bank of the Danube River.

astro-ph.CO

The precision of slow-roll predictions for the CMBR anisotropies

Inflationary predictions for the anisotropy of the cosmic microwave background radiation (CMBR) are often based on the slow-roll approximation. We study the precision with which the multipole moments of the temperature two-point correlation function can be predicted by means of the slow-roll approximation. We ask whether this precision is good enough for the forthcoming high precision observations by means of the MAP and Planck satellites. The error in the multipole moments due to the slow-roll approximation is demonstrated to be bigger than the error in the power spectrum. For power-law inflation with $n_S=0.9$ the error from the leading order slow-roll approximation is $\approx 5%$ for the amplitudes and $\approx 20%$ for the quadrupoles. For the next-to-leading order the errors are within a few percent. The errors increase with $|n_S - 1|$. To obtain a precision of 1% it is necessary, but in general not sufficient, to use the next-to-leading order. In the case of power-law inflation this precision is obtained for the spectral indices if $|n_S - 1| < 0.02$ and for the quadrupoles if $|n_S - 1| < 0.15$ only. The errors in the higher multipoles are even larger than those for the quadrupole, e.g. $\approx 15%$ for l=100, with $n_S = 0.9$ at the next-to-leading order. We find that the accuracy of the slow-roll approximation may be improved by shifting the pivot scale of the primordial spectrum (the scale at which the slow-roll parameters are fixed) into the regime of acoustic oscillations. Nevertheless, the slow-roll approximation cannot be improved beyond the next-to-leading order in the slow-roll parameters.

astro-ph