SearcharxivSearch

arXiv subjects

Sean Richardson

Publications and source records attributed to Sean Richardson.

8 recordsLinked to original sources

Marked magnetic action rigidity

An exact magnetic system over a closed manifold $M$ consists of a pair $(g,\alpha)$, where $g$ is a Riemannian metric and $\alpha$ is a 1-form encoding a magnetic field. In this context, we consider a generalization of the marked length rigidity conjecture: does the marked magnetic action spectrum of magnetic systems with Anosov magnetic flow determine the metric and the 1-form, up to a natural obstruction? In this article we answer this question in two settings: 1) locally for systems with close metrics and 1-forms and 2) for metrics in the same conformal class.

math.DS

Guillarmou's Normal Operator for Magnetic and Thermostat Flows

Guillarmou's normal operator over a closed Anosov manifold is analogous to the classical normal operator of the geodesic X-ray transform over manifolds with boundary. In this paper, we generalize this normal operator, under some dynamical assumptions, to thermostat flows as well as to the case of the magnetic flows. In particular, we show that these generalized normal operators are elliptic pseudodifferential operators of order -1 in each case. As an application, we prove a stability estimate for the magnetic X-ray transform.

math.AP

Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers

When an LLM learns a new fact during finetuning (e.g., new movie releases, newly elected pope, etc.), where does this information go? Are entities enriched with relation information immediately, or do models recall information just-in-time before a prediction? Or, are "all of the above" true, with LLMs implementing multiple redundant heuristics? Existing localization approaches (e.g., activation patching) are ill-suited for this analysis because they usually replace parts of the residual stream, thus overriding previous information. To fill this interpretability gap, we propose dynamic weight grafting, an analysis technique that selectively grafts subsets of weights from a finetuned model onto a pretrained model. Using this technique, we show two separate pathways for retrieving finetuned relation information: 1) "enriching" the residual stream with relation information while processing the tokens that correspond to an entity (e.g., "Zendaya" in "Zendaya co-starred with Timoth\'ee Chalamet" and 2) "recalling" this information at the final token position before generating a target fact. In some cases, models need information from both of these pathways to correctly generate finetuned facts while, in other cases, either the "enrichment" or "recall" pathway alone is sufficient. We localize the "recall" pathway to model components -- finding that "recall" occurs via both task-specific attention mechanisms and an entity-specific extraction step in the feedforward networks of the final layers before prediction. By targeting model components and parameters, as opposed to just activations, we are able to understand the mechanisms by which finetuned knowledge is retrieved during generation.

cs.LG

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.

cs.CL

An inversion formula for the X-ray normal operator over closed hyperbolic surfaces

We construct an explicit inversion formula for Guillarmou's normal operator on closed surfaces of constant negative curvature. This normal operator can be defined as a weak limit for an "attenuated normal operator", and we prove this inversion formula by first constructing an additional inversion formula for this attenuated normal operator on both the Poincar\'e disk and closed surfaces of constant negative curvature. A consequence of the inversion formula is the explicit construction of invariant distributions with prescribed pushforward over closed hyperbolic manifolds.

math.DG

RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective method for measuring the sensitivity of a reward model to high-level attributes of responses, such as sentiment, helpfulness, or complexity. Importantly, RATE measures the causal effect of an attribute on the reward. RATE uses LLMs to rewrite responses to produce imperfect counterfactuals examples that can be used to measure causal effects. A key challenge is that these rewrites are imperfect in a manner that can induce substantial bias in the estimated sensitivity of the reward model to the attribute. The core idea of RATE is to adjust for this imperfect-rewrite effect by rewriting twice. We establish the validity of the RATE procedure and show empirically that it is an effective estimator.

cs.CL

A Sharp Fourier Inequality and the Epanechnikov Kernel

We consider functions $f: \mathbb{Z} \to \mathbb{R}$ and kernels $u: \{-n, \cdots, n\} \to \mathbb{R}$ normalized by $\sum_{\ell = -n}^{n} u(\ell) = 1$, making the convolution $u \ast f$ a "smoother" local average of $f$. We identify which choice of $u$ most effectively smooths the second derivative in the following sense. For each $u$, basic Fourier analysis implies there is a constant $C(u)$ so $\|Δ(u \ast f)\|_{\ell^2(\mathbb{Z})} \leq C(u)\|f\|_{\ell^2(\mathbb{Z})}$ for all $f: \mathbb{Z} \to \mathbb{R}$. By compactness, there is some $u$ that minimizes $C(u)$ and in this paper, we find explicit expressions for both this minimal $C(u)$ and the minimizing kernel $u$ for every $n$. The minimizing kernel is remarkably close to the Epanechnikov kernel in Statistics. This solves a problem of Kravitz-Steinerberger and an extremal problem for polynomials is solved as a byproduct.

math.CA

You can hear the local orientability of an orbifold

A Riemannian orbifold is a mildly singular generalization of a Riemannian manifold which is locally modeled on the quotient of a connected, open manifold under a finite group of isometries. If all of the isometries used to define the local structures of an entire orbifold are orientation preserving, we call the orbifold locally orientable. We use heat invariants to show that a Riemannian orbifold which is locally orientable cannot be Laplace isospectral to a Riemannian orbifold which is not locally orientable. As a corollary we observe that a Riemannian orbifold that is not locally orientable cannot be Laplace isospectral to a Riemannian manifold.

math.DG