SearcharxivSearch

arXiv subjects

Bernie Hogan

Publications and source records attributed to Bernie Hogan.

4 recordsLinked to original sources

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.

cs.CL

Towards a Harms Taxonomy of AI Likeness Generation

Generative artificial intelligence models, when trained on a sufficient number of a person's images, can replicate their identifying features in a photorealistic manner. We refer to this process as 'likeness generation'. Likeness-featuring synthetic outputs often present a person's likeness without their control or consent, and may lead to harmful consequences. This paper explores philosophical and policy issues surrounding generated likeness. It begins by offering a conceptual framework for understanding likeness generation by examining the novel capabilities introduced by generative systems. The paper then establishes a definition of likeness by tracing its historical development in legal literature. Building on this foundation, we present a taxonomy of harms associated with generated likeness, derived from a comprehensive meta-analysis of relevant literature. This taxonomy categorises harms into seven distinct groups, unified by shared characteristics. Utilising this taxonomy, we raise various considerations that need to be addressed for the deployment of appropriate mitigations. Given the multitude of stakeholders involved in both the creation and distribution of likeness, we introduce concepts such as indexical sufficiency, a distinction between generation and distribution, and harms as having a context-specific nature. This work aims to serve industry, policymakers, and future academic researchers in their efforts to address the societal challenges posed by likeness generation.

cs.CY

A Time Decoupling Approach for Studying Forum Dynamics

Online forums are rich sources of information about user communication activity over time. Finding temporal patterns in online forum communication threads can advance our understanding of the dynamics of conversations. The main challenge of temporal analysis in this context is the complexity of forum data. There can be thousands of interacting users, who can be numerically described in many different ways. Moreover, user characteristics can evolve over time. We propose an approach that decouples temporal information about users into sequences of user events and inter-event times. We develop a new feature space to represent the event sequences as paths, and we model the distribution of the inter-event times. We study over 30,000 users across four Internet forums, and discover novel patterns in user communication. We find that users tend to exhibit consistency over time. Furthermore, in our feature space, we observe regions that represent unlikely user behaviors. Finally, we show how to derive a numerical representation for each forum, and we then use this representation to derive a novel clustering of multiple forums.

cs.SI

Modeling the evolution of continuously-observed networks: Communication in a Facebook-like community

Building on existing stochastic actor-oriented models for panel data, we employ a conditional logistic framework to explore growth mechanisms for tie creation in continuously-observed networks. This framework models the likelihood of tie formation distinguishing it from hazard models that consider time to tie formation. It enables multiple growth mechanisms for network evolution (homophily, focus constraints, reinforcement, reciprocity, triadic closure, and popularity) to be modeled simultaneously. We apply this framework to communication within a Facebook-like community. The findings exemplify the inadequacy of descriptive measures that test single mechanisms independently. They also indicate how system design shapes behavior and network evolution.

physics.soc-ph