SearcharxivSearch

arXiv subjects

Raphael Sarfati

Publications and source records attributed to Raphael Sarfati.

11 recordsLinked to original sources

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. This abstraction gives practitioners little visibility into what their data actually teaches models, allowing spurious correlations to be learned by a model and inducing undesirable behaviors such as over-stylization and sycophancy. To address this problem, we ask: can we inspect a preference dataset before optimization and decide, at the level of concepts, which behaviors a model should be allowed to learn? Motivated by this, we introduce a data-centric post-training pipeline that uses interpretability protocols to develop statistical hypotheses for the latent concepts separating preferred from dispreferred generations, making them explicit for fine-grained user feedback. Building on this view, we unify several interpretability-based training protocols as ways of shaping rewards via feature or data interventions. Empirically, we show that our pipeline diagnoses undesirable signals in existing preference data, mitigates off-target learning, and can also help amplify or shape desired properties such as safeguards and model personality. More broadly, our results suggest that interpretability can turn post-training from optimizing opaque proxy rewards into a process of auditing and sculpting the learning signal itself.

cs.LG

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

Neural representations carry rich geometric structure; but does that structure causally shape behavior? To address this question, we intervene along paths through activation space defined by different geometries, and measure the behavioral trajectories they induce. In particular, we test whether interventions that respect the geometry of activation space will yield behaviors close to those the model exhibits naturally. Concretely, we first fit an activation manifold $M_h$ to representations and a behavior manifold $M_y$ to output probability distributions. We then test the link $M_h \leftrightarrow M_y$ via interventions: we find that steering along $M_h$, which we term manifold steering, yields behavioral trajectories that follow $M_y$, while linear steering -- which assumes a Euclidean geometry -- cuts through off-manifold regions and hence produces unnatural outputs. Moreover, optimizing interventions in activation space to produce paths along $M_y$ recovers activation trajectories that trace the curvature of $M_h$. We demonstrate this bidirectional relationship between the geometry of representation and behavior across tasks and modalities. In language models, we use reasoning tasks with cyclic and sequential geometries as well as in-context learning tasks with more complex graph geometries. In a video world model, we use a task with geometry corresponding to physical dynamics. Overall, our work shows that geometry in neural representation is not merely incidental, but is in fact the proper object for enabling principled control via intervention on internals. This recasts the core problem of steering from finding the right direction to finding the right geometry.

cs.LG

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer, but continues generating tokens without revealing its internal belief. Our analysis compares activation probing, early forced answering, and a CoT monitor across two large models (DeepSeek-R1 671B & GPT-OSS 120B) and find task difficulty-specific differences: The model's final answer is decodable from activations far earlier in CoT than a monitor is able to say, especially for easy recall-based MMLU questions. We contrast this with genuine reasoning in difficult multihop GPQA-Diamond questions. Despite this, inflection points (e.g., backtracking, 'aha' moments) occur almost exclusively in responses where probes show large belief shifts, suggesting these behaviors track genuine uncertainty rather than learned "reasoning theater." Finally, probe-guided early exit reduces tokens by up to 80% on MMLU and 30% on GPQA-Diamond with similar accuracy, positioning attention probing as an efficient tool for detecting performative reasoning and enabling adaptive computation.

cs.CL

Tracking and triangulating firefly flashes in field recordings

Identifying firefly flashes from other bright features in nature images is complicated. I provide a training dataset and trained neural networks for reliable flash classification. The training set consists of thousands of cropped images (patches) extracted by manual labeling from video recordings of fireflies in their natural habitat. The trained network appears as considerably more reliable to differentiate flashes from other sources of light compared to traditional methods relying solely on intensity thresholding. This robust tracking enables a new calibration-free method for the 3D reconstruction of flash occurrences from stereoscopic 360-degree videos, which I also present here.

cs.CV

Unravelling the origins of anomalous diffusion: from molecules to migrating storks

Anomalous diffusion or, more generally, anomalous transport, with nonlinear dependence of the mean-squared displacement on the measurement time, is ubiquitous in nature. It has been observed in processes ranging from microscopic movement of molecules to macroscopic, large-scale paths of migrating birds. Using data from multiple empirical systems, spanning 12 orders of magnitude in length and 8 orders of magnitude in time, we employ a method to detect the individual underlying origins of anomalous diffusion and transport in the data. This method decomposes anomalous transport into three primary effects: long-range correlations ("Joseph effect"), fat-tailed probability density of increments ("Noah effect"), and non-stationarity ("Moses effect"). We show that such a decomposition of real-life data allows to infer nontrivial behavioral predictions, and to resolve open questions in the fields of single particle tracking in living cells and movement ecology.

physics.data-an

Hydrodynamics of a dense flock of sheep: edge motion and long-range correlations

Sheep are gregarious animals, and they often aggregate into dense, cohesive flocks, especially under stress. In this paper, we use image processing tools to analyze a publicly available aerial video showing a dense sheep flock moving under the stimulus of a shepherding dog. Inspired by the fluidity of the motion, we implement a hydrodynamics approach, extracting velocity fields, and measuring their propagation and correlations in space and time. We find that while the flock overall is stationary, significant dynamics happens at the edges, notably in the form of fluctuations propagating like waves, and large-scale correlations spanning the entire flock. These observations highlight the importance of incorporating interfacial dynamics, for instance in the form of line tension, when using a hydrodynamics framework to model the dynamics of dense, non-polarized swarms.

physics.bio-ph

Temporally Anticorrelated Subdiffusion in Water Nanofilms on Silica Suggests Near-Surface Viscoelasticity

We used single-molecule tracking to probe the local rheology of interfacial water. Fluorescent rhodamine molecules were tracked on silica surfaces as a function of ambient relative humidity, which controlled the thickness of condensed water nanofilms. At low humidity, the molecules exhibited confined diffusion in the vicinity of isolated adsorption sites characterized by a broad distribution of binding stiffness constants; subsequent chemical or physical surface passivation selectively eliminated stiffer binding sites. At increased humidity, molecularly thin water films condensed, permitting near-surface transport of rhodamine molecules. Motion was subdiffusive, with an anomalous exponent increasing with the nanofilm thickness. Molecular trajectories were temporally anticorrelated, ergodic, but also featured transient binding and intermittent diffusion. Statistical modeling demonstrated that this complex motion in water nanofilms had the characteristics of fractional Brownian motion combined with a continuous time random walk. This was consistent with diffusion within viscoelastic nanofilms, suggesting persistent molecular structuring in the vicinity of the silica surface.

cond-mat.soft

Long-range attraction of particles adhered to lipid vesicles

Many biological systems fold thin sheets of lipid membrane into complex three-dimensional structures. This microscopic origami is often mediated by the adsorption and self-assembly of proteins on a membrane. As a model system to study adsorption-mediated interactions, we study the collective behavior of micrometric particles adhered to a lipid vesicle. We estimate the colloidal interactions using a maximum likelihood analysis of particle trajectories. When the particles are highly wrapped by a tense membrane, we observe strong long-range attractions with a typical binding energy of 150 $k_B T$ and significant forces extending a few microns.

cond-mat.soft

Mechanical stability of particle-stabilized droplets under micropipette aspiration

We investigate the mechanical behavior of particle-stabilized droplets using micropipette aspiration. We observe that droplets stabilized with amphiphilic dumbbell-shaped particles exhibit a two-stage response to increasing suction pressure. Droplets first drip, then wrinkle and buckle like an elastic shell. While particles have a dramatic impact on the mechanism of failure, the mechanical strength of the droplets is only modestly increased. On the other hand, droplets coated with the molecular surfactant Sodium Dodecyl Sulfate are even weaker than bare droplets. In all cases, the magnitude of the critical pressure for the onset of instabilities is set by the fluid surface tension.

cond-mat.soft

Maximum likelihood estimations of force and mobility from short single Brownian trajectories

We describe a method to extract force and diffusion parameters from single trajectories of Brownian particles based on the principle of maximum likelihood. The analysis is well-suited for out-of-equilibrium trajectories, even when a limited amount of data is available and the dynamical parameters vary spatially. We substantiate this method with experimental and simulated data, and discuss its practical implementation, strengths, and limitations.

cond-mat.soft