SearcharxivSearch

arXiv subjects

Eduard Kapelko

Publications and source records attributed to Eduard Kapelko.

3 recordsLinked to original sources

A Pigouvian Matchmaker Mechanism for De-escalating the AGI Race

A formal mechanism is presented in which a willing regulator-matchmaker fosters cooperation on resources among participants in the AGI race, collects a Pigouvian tax based on the speed-up it induces, and invests the proceeds into alignment research. The construction is derived in the continuous-time options framework of Tan (2025) in which cooperation is treated as a jump in the underlying asset value of participating players, the Pigouvian component is matched to the marginal effect of increasing expected loss, and the total collected fund endogenizes the rate of learning on safety. It is shown how the framework allows for determining participation and optimal activity levels. Conditions under which it is optimal to enter the market are derived, and it is proven that if the orthogonality condition holds between the supported portfolio and the abilities component, the Suicide Region collapses at finite time, and the upper bound for this time is derived as sum of a deterministic and random term. Finally, if orthogonality is violated, it is proven that enhancing matchmaker capacity does not recover the market's superiority. The construction links research areas including two-sided markets, Pigouvian taxes, self-regulatory organizations, private law enforcement, evolutionary modeling of AI races, real options and option games, measurement of comparative progress and analysis of the Suicide Region.

cs.GT

Active Inference Agency Formalization, Metrics, and Convergence Assessments

This paper addresses the critical challenge of mesa-optimization in AI safety by providing a formal definition of agency and a framework for its analysis. Agency is conceptualized as a Continuous Representation of accumulated experience that achieves autopoiesis through a dynamic balance between curiosity (minimizing prediction error to ensure non-computability and novelty) and empowerment (maximizing the control channel's information capacity to ensure subjectivity and goal-directedness). Empirical evidence suggests that this active inference-based model successfully accounts for classical instrumental goals, such as self-preservation and resource acquisition. The analysis demonstrates that the proposed agency function is smooth and convex, possessing favorable properties for optimization. While agentic functions occupy a vanishingly small fraction of the total abstract function space, they exhibit logarithmic convergence in sparse environments. This suggests a high probability for the spontaneous emergence of agency during the training of modern, large-scale models. To quantify the degree of agency, the paper introduces a metric based on the distance between the behavioral equivalents of a given system and an "ideal" agentic function within the space of canonicalized rewards (STARC). This formalization provides a concrete apparatus for classifying and detecting mesa-optimizers by measuring their proximity to an ideal agentic objective, offering a robust tool for analyzing and identifying undesirable inner optimization in complex AI systems.

cs.LG

Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI

Safety and controllability are critical for large language models. A central question is whether undesirable behaviors like deception are localized functions that can be removed, or if they are deeply intertwined with a model's core cognitive abilities. We introduce "cyclic ablation," an iterative method to test this. By combining sparse autoencoders, targeted ablation, and adversarial training on DistilGPT-2, we attempted to eliminate the concept of deception. We found that, contrary to the localization hypothesis, deception was highly resilient. The model consistently recovered its deceptive behavior after each ablation cycle via adversarial training, a process we term functional regeneration. Crucially, every attempt at this "neurosurgery" caused a gradual but measurable decay in general linguistic performance, reflected by a consistent rise in perplexity. These findings are consistent with the view that complex concepts are distributed and entangled, underscoring the limitations of direct model editing through mechanistic interpretability.

cs.CL