SearcharxivSearch

arXiv subjects

Aleksandr Bowkis

Publications and source records attributed to Aleksandr Bowkis.

2 recordsLinked to original sources

Automated alignment is harder than you think

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically flawed). Consequently, research outputs will contain systematic, undetected errors, and even correct outputs could be incorrectly aggregated into overconfident safety assessments. This problem is likely to be worse for automated alignment research than for human-generated alignment research for several reasons: 1) optimisation pressure means agent-generated mistakes are concentrated among those that human reviewers are least likely to catch; 2) agents are likely to produce errors that do not resemble human mistakes; 3) AI-generated alignment solutions may involve arguments humans cannot evaluate; and 4) shared weights, data and training processes may make AI outputs more correlated than human equivalents. Therefore, agents must be trained to reliably perform hard-to-supervise fuzzy tasks. Generalisation and scalable oversight are the leading candidates for achieving this but both face novel challenges in the context of automated alignment.

cs.AI

The reconstructed CMB lensing bispectrum

Weak gravitational lensing by the intervening large-scale structure (LSS) of the Universe is the leading non-linear effect on the anisotropies of the cosmic microwave background (CMB). The integrated line-of-sight mass that causes the distortion -- known as lensing convergence -- can be reconstructed from the lensed temperature and polarization anisotropies via estimators quadratic in the CMB modes, and its power spectrum has been measured from multiple CMB experiments. Sourced by the non-linear evolution of structure, the bispectrum of the lensing convergence provides additional information on late-time cosmological evolution complementary to the power spectrum. However, when trying to estimate the summary statistics of the reconstructed lensing convergence, a number of noise-biases are introduced, as previous studies have shown for the power spectrum. Here, we explore for the first time the noise-biases in measuring the bispectrum of the reconstructed lensing convergence. We compute the leading noise-biases in the flat-sky limit and compare our analytical results against simulations, finding excellent agreement. Our results are critical for future attempts to reconstruct the lensing convergence bispectrum with real CMB data.

astro-ph.CO