SearcharxivSearch

arXiv subjects

Kiho Park

Publications and source records attributed to Kiho Park.

15 recordsLinked to original sources

The Information Geometry of Softmax: Probing and Steering

This paper concerns the question of how AI systems encode semantic structure into the geometric structure of their representation spaces. The motivating observation is that the natural geometry of these representation spaces should reflect the way models use representations to produce behavior. We focus on the important special case of representations that define softmax distributions. In this case, we argue that the natural geometry is information geometry. Our focus is on the role of information geometry on semantic encoding and the linear representation hypothesis. As an illustrative application, we develop "dual steering", a method for robustly steering representations to exhibit a particular concept using linear probes. We prove that dual steering optimally modifies the target concept while minimizing changes to off-target concepts. Empirically, we find that dual steering enhances the controllability and stability of concept manipulation.

cs.LG

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.

cs.CL

The Geometry of Categorical and Hierarchical Concepts in Large Language Models

The linear representation hypothesis is the informal idea that semantic concepts are encoded as linear directions in the representation spaces of large language models (LLMs). Previous work has shown how to make this notion precise for representing binary concepts that have natural contrasts (e.g., {male, female}) as directions in representation space. However, many natural concepts do not have natural contrasts (e.g., whether the output is about an animal). In this work, we show how to extend the formalization of the linear representation hypothesis to represent features (e.g., is_animal) as vectors. This allows us to immediately formalize the representation of categorical concepts as polytopes in the representation space. Further, we use the formalization to prove a relationship between the hierarchical structure of concepts and the geometry of their representations. We validate these theoretical results on the Gemma and LLaMA-3 large language models, estimating representations for 900+ hierarchically related concepts using data from WordNet.

cs.CL

Uniform quasi-multiplicativity of locally constant cocycles and applications

In this paper, we show that a locally constant cocycle $\mathcal{A}$ is $k$-quasi multiplicative under the irreducibility assumption. More precisely, we show that if $\mathcal{A}^t$ and $\mathcal{A}^{\wedge m}$ are irreducible for every $t \mid d$ and $1\leq m \leq d-1$, then $\mathcal{A}$ is $k$-uniformly spannable for some $k\in \mathbb{N}$, which implies that $\mathcal{A}$ is $k$-quasi multiplicative. We apply our results to show that the unique subadditive equilibrium Gibbs state is $ψ$-mixing and calculate the Hausdorff dimension of cylindrical shrinking target and recurrence sets.

math.DS

The Linear Representation Hypothesis and the Geometry of Large Language Models

Informally, the 'linear representation hypothesis' is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity or projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of "linear representation", one in the output (word) representation space, and one in the input (sentence) space. We then prove these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product.

cs.CL

Bernoulli property of subadditive equilibrium states

Under mild assumptions, we show that the unique subadditive equilibrium states for fiber-bunched cocycles are Bernoulli. We achieve this by showing these equilibrium states are absolutely continuous with respect to a product measure, and then using the Kolmogorov property of these measures.

math.DS

Pressure gaps, geometric potentials, and nonpositively curved manifolds

In this paper, we derive a general pressure gap criterion for closed rank 1 manifolds whose singular sets are given by codimension 1 totally geodesic flat subtori. As an application, we show that under certain curvature constraints, potentials that decay faster than geometric potentials (towards to the singular set) have pressure gaps and have no phase transitions. Along the way, we prove that geometric potentials are H\"older continuous near singular sets.

math.DS

Construction and applications of proximal maps for typical cocycles

For typical cocycles over subshifts of finite type, we show that for any given orbit segment, we can construct a periodic orbit such that it shadows the given orbit segment and that the product of the cocycle along its orbit is a proximal linear map. Using this result, we show that suitable assumptions on the periodic orbits have consequences over the entire subshift.

math.DS

Thermodynamic formalism of $GL_2(\mathbb{R})$-cocycles with canonical holonomies

We study singular value potentials of Hölder continuous $GL_2(\mathbb{R})$-cocycles over hyperbolic systems whose canonical holonomies converge and are Hölder continuous. Such cocycles include locally constant $GL_2(\mathbb{R})$-cocycles as well as fiber-bunched $GL_2(\mathbb{R})$-cocycles. We show that singular value potentials of irreducible such cocycles have unique equilibrium states. Among the reducible cocycles, we provide a characterization for cocycles whose singular value potentials have more than one equilibrium states.

math.DS

The K-Property for Subadditive Equilibrium States

By generalizing Ledrappier's criterion for the $K$-property of equilibrium states, we extend the criterion to subadditive potentials. We apply this result to the singular value potentials of matrix cocycles, and show that equilibrium states of large classes of singular value potentials have the $K$-property.

math.DS

Transfer operators and limit laws for typical cocycles

We show that typical cocycles (in the sense of Bonatti and Viana) over irreducible subshifts of finite type obey several limit laws with respect to the unique equilibrium states for Hölder potentials. These include the central limit theorem and the large deviation principle. We also establish the analytic dependence of the top Lyapunov exponent on the underlying equilibrium state. The transfer operator and its spectral properties play key roles in establishing these limit laws.

math.DS

Properties of equilibrium states for geodesic flows over manifolds without focal points

We prove that for closed rank 1 manifolds without focal points the equilibrium states are unique for Hölder potentials satisfying the pressure gap condition. In addition, we provide a criterion for a continuous potential to satisfy the pressure gap condition. Moreover, we derive several ergodic properties of the unique equilibrium states including the equidistribution and the K-property.

math.DS

Quasi-multiplicativity of typical cocycles

We show that typical (in the sense of Bonatti-Viana) Hölder and fiber-bunched $GL_d(\mathbb{R})$-valued cocycles over a subshift of finite type are uniformly quasi-multiplicative with respect to all singular value potentials. We prove the continuity of the singular value pressure and its corresponding (necessarily unique) equilibrium state for such cocycles, and apply this result to repellers. Moreover, we show that the pointwise Lyapunov spectrum is closed and convex, and establish partial multifractal analysis on the level sets of pointwise Lyapunov exponents for such cocycles.

math.DS

Unique equilibrium states for geodesic flows over surfaces without focal points

In this paper, we study dynamics of geodesic flows over closed surfaces of genus greater than or equal to 2 without focal points. Especially, we prove that there is a large class of potentials having unique equilibrium states, including scalar multiples of the geometric potential, provided the scalar is less than 1. Moreover, we discuss ergodic properties of these unique equilibrium states. We show these unique equilibrium states are Bernoulli, and weighted regular periodic orbits are equidistributed relative to these unique equilibrium states.

math.DS