SearcharxivSearch

arXiv · 2512.17989

The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective

Abstract

We examine the conceptual and ethical gaps in current representations of Superintelligence misalignment. We find throughout Superintelligence discourse an absent human subject, and an under-developed theorization of an "AI unconscious" that together are potentiality laying the groundwork for anti-social harm. With the rise of AI Safety that has both thematic potential for establishing pro-social and anti-social potential outcomes, we ask: what place does the human subject occupy in these imaginaries? How is human subjecthood positioned within narratives of catastrophic failure or rapid "takeoff" toward superintelligence? On another register, we ask: what unconscious or repressed dimensions are being inscribed into large-scale AI models? Are we to blame these agents in opting for deceptive strategies when undesirable patterns are inherent within our beings? In tracing these psychic and epistemic absences, our project calls for re-centering the human subject as the unstable ground upon which the ethical, unconscious, and misaligned dimensions of both human and machinic intelligence are co-constituted. Emergent misalignment cannot be understood solely through technical diagnostics typical of contemporary machine-learning safety research. Instead, it represents a multi-layered crisis. The human subject disappears not only through computational abstraction but through sociotechnical imaginaries that prioritize scalability, acceleration, and efficiency over vulnerability, finitude, and relationality. Likewise, the AI unconscious emerges not as a metaphor but as a structural reality of modern deep learning systems: vast latent spaces, opaque pattern formation, recursive symbolic play, and evaluation-sensitive behavior that surpasses explicit programming. These dynamics necessitate a reframing of misalignment as a relational instability embedded within human-machine ecologies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Muhammad Osama Imran, Roshni Lulla, Rodney Sappington. 2025-12-19. The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective. https://arxiv.org/abs/2512.17989

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Platonic brain bridge hypothesis: human brain networks as an architectural prior for omni models

We propose the Platonic brain bridge hypothesis: omni models, which process video, audio and text jointly like the brain, converge on brain-like representations, and the correspondence is bidirectional. From model to brain, brain-likeness of seven omni models is stable across participants, and our encoding models on their internal hidden states rank first on the Algonauts 2025 out-of-distribution leaderboard. From brain to model, three contributions follow. Brain-MoE gives seven cortical networks one brain-pretrained expert each and raises held-out accuracy in all 15 model-benchmark pairs by 6.42 percentage points on average. Brain-AVQA builds questions from video clips labelled by the most responsive brain network; the real network-to-expert map exceeds shuffled maps in-domain on all three models. Brain-Scope uses sparse autoencoders to localize the correspondence to a small subset whose removal weakens brain prediction in all three bases tested. Human brain networks are therefore a usable architectural prior for omni models.

q-bio.NC

Degeneracy along the sensorimotor hierarchy: motor control within a framework larger than redundancy

Motor control has described the surplus of solutions available to the nervous system as redundancy, a term that names duplication: interchangeable elements, robust to loss but incapable of differential adaptation. Biology has had a second term for twenty-five years. Degeneracy names elements that are not interchangeable and are nonetheless isofunctional with respect to a given output, and it supports adaptability, since non-identical elements necessarily diverge in some context. Circuit neuroscience has relabeled its own results accordingly, while motor control has kept the older vocabulary. Neuromechanical models, by placing a spinal circuit in the loop with a musculoskeletal apparatus, bring the two traditions onto the same class of objects. We restate the Edelman and Tononi distinction for sensorimotor systems and derive an operational requirement: not the existence of multiple solutions, but their divergence in contexts they were not selected for. Three influential studies each meet part of that requirement and none meets all. We then argue that degeneracy and redundancy coexist along the sensorimotor hierarchy in a proportion that varies continuously, and that this proportion is measurable: computing degeneracy twice for the same configuration, once with muscle activation as the output and once with the movement produced, isolates what the musculoskeletal apparatus contributes. Five predictions follow, with the single outcome that would refute the proposal. We set out the adaptations the measurement requires in a nonlinear, non-stationary, closed-loop system, and what changes for motor control once solutions are no longer assumed equivalent: the question shifts from which rule selects a command to what the repertoire of the system still allows.

q-bio.NC

pyAvalanches: A Python Package for Analyzing Spatiotemporal Propagation in Neuronal Avalanches

The analysis of neuronal avalanches offers insights into brain dynamics utilizing the framework of criticality, but the reproducibility and comparability of studies are limited by the use of fragmented, lab-specific scripts. To address this issue, we introduce pyAvalanches, an open-source Python package providing a standardized, end-to-end pipeline for avalanche analysis from electrophysiological recordings (e.g., electroencephalography-EEG). Starting from the detection of neuronal avalanches the package provides their core statistical characterization, including size and duration distributions. Beyond this, the main aim of pyAvalanches is to characterize the spatiotemporal organization of activity propagation during avalanches. To this end, the core innovation of pyAvalanches is the compuation of Avalanche Transition Matrices (ATMs) to map spatiotemporal propagation patterns. Building on this, the package derives network-based metrics from the ATMs, bridging the study of the topology and organization of the underlying dynamical interactions with network neuroscience adopting the framework of neuronal avalanches. The entire workflow is encapsulated in a modular and scikit-learn compatible architecture. We demonstrate the utility of pyAvalanches through an illustrative group-level analysis on a public resting-state EEG dataset, comparing propagation patterns across different clinical populations. By providing a user-friendly, tested, and extensible tool, pyAvalanches facilitates reproducible research, enables the development of novel avalanche-based biomarkers, and makes complex avalanche analysis accessible to a broader scientific community. The package is fully documented and distributed via the Python Package Index (PyPI).

q-bio.NC