Searcharxiv⌕ Search

arXiv subjects

Michael Hahn

Publications and source records attributed to Michael Hahn.

At least 37 records · Page 2Linked to original sources

The Expressive Capacity of State Space Models: A Formal Language Perspective

Recently, recurrent models based on linear state space models (SSMs) have shown promising performance in language modeling (LM), competititve with transformers. However, there is little understanding of the in-principle abilities of such models, which could provide useful guidance to the search for better LM architectures. We present a comprehensive theoretical study of the capacity of such SSMs as it compares to that of transformers and traditional RNNs. We find that SSMs and transformers have overlapping but distinct strengths. In star-free state tracking, SSMs implement straightforward and exact solutions to problems that transformers struggle to represent exactly. They can also model bounded hierarchical structure with optimal memory even without simulating a stack. On the other hand, we identify a design choice in current SSMs that limits their expressive power. We discuss implications for SSM and LM research, and verify results empirically on a recent SSM, Mamba.

cs.CL↗

Softmax Transformers are Turing-Complete

Hard attention Chain-of-Thought (CoT) transformers are known to be Turing-complete. However, it is an open problem whether softmax attention Chain-of-Thought (CoT) transformers are Turing-complete. In this paper, we prove a stronger result that length-generalizable softmax CoT transformers are Turing-complete. More precisely, our Turing-completeness proof goes via the CoT extension of the Counting RASP (C-RASP), which correspond to softmax CoT transformers that admit length generalization. We prove Turing-completeness for CoT C-RASP with causal masking over a unary alphabet (more generally, for letter-bounded languages). While we show this is not Turing-complete for arbitrary languages, we prove that its extension with relative positional encoding is Turing-complete for arbitrary languages. We empirically validate our theory by training transformers for languages requiring complex (non-linear) arithmetic reasoning.

cs.FL↗

Linguistic Structure from a Bottleneck on Sequential Information Processing

Human language has a distinct systematic structure, where utterances break into individually meaningful words which are combined to form phrases. We show that natural-language-like systematicity arises in codes that are constrained by a statistical measure of complexity called predictive information, also known as excess entropy. Predictive information is the mutual information between the past and future of a stochastic process. In simulations, we find that such codes break messages into groups of approximately independent features which are expressed systematically and locally, corresponding to words and phrases. Next, drawing on crosslinguistic text corpora, we find that actual human languages are structured in a way that reduces predictive information compared to baselines at the levels of phonology, morphology, syntax, and lexical semantics. Our results establish a link between the statistical and algebraic structure of language and reinforce the idea that these structures are shaped by communication under general cognitive constraints.

cs.CL↗

Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities

Transformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in large-scale pretrained LLMs, or whether LLMs might effectively overcome these constraints in practice due to the scale of both the models themselves and their pretraining data. We explore how these architectural constraints manifest after pretraining, by studying a family of $\textit{retrieval}$ and $\textit{copying}$ tasks inspired by Liu et al. [2024a]. We use a recently proposed framework for studying length generalization [Huang et al., 2025] to provide guarantees for each of our settings. Empirically, we observe an $\textit{induction-versus-anti-induction}$ asymmetry, where pretrained models are better at retrieving tokens to the right (induction) rather than the left (anti-induction) of a query token. This asymmetry disappears upon targeted fine-tuning if length-generalization is guaranteed by theory. Mechanistic analysis reveals that this asymmetry is connected to the differences in the strength of induction versus anti-induction circuits within pretrained transformers. We validate our findings through practical experiments on real-world tasks demonstrating reliability risks. Our results highlight that pretraining selectively enhances certain transformer capabilities, but does not overcome fundamental length-generalization limits.

cs.LG↗

Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B

In-Context Learning (ICL) is an intriguing ability of large language models (LLMs). Despite a substantial amount of work on its behavioral aspects and how it emerges in miniature setups, it remains unclear which mechanism assembles task information from the individual examples in a fewshot prompt. We use causal interventions to identify information flow in Gemma-2 2B for five naturalistic ICL tasks. We find that the model infers task information using a two-step strategy we call contextualize-then-aggregate: In the lower layers, the model builds up representations of individual fewshot examples, which are contextualized by preceding examples through connections between fewshot input and output tokens across the sequence. In the higher layers, these representations are aggregated to identify the task and prepare prediction of the next output. The importance of the contextualization step differs between tasks, and it may become more important in the presence of ambiguous examples. Overall, by providing rigorous causal analysis, our results shed light on the mechanisms through which ICL happens in language models.

cs.CL↗

Formation of a Coronal Hole by a quiet-Sun Filament Eruption

A coronal hole formed as a result of a quiet-Sun filament eruption close to the solar disk center on 2014 June 25. We studied this formation using images from the Atmospheric Imaging Assembly (AIA), magnetograms from the Helioseismic and Magnetic Imager (HMI), and a differential emission measure (DEM) analysis derived from the AIA images. The coronal hole developed in three stages: (1) formation, (2) migration, and (3) stabilization. In the formation phase, the emission measure (EM) and temperature started to decrease six hours before the filament erupted. Then, the filament erupted and a large coronal dimming formed over the following three hours. Subsequently, in a phase lasting $15.5$~hours, the coronal dimming migrated by 150" from its formation site to a location where potential field source surface extrapolations indicate the presence of open magnetic field lines, marking the transition into a coronal hole. During this migration, the coronal hole drifted across quasi-stationary magnetic elements in the photosphere, implying the occurrence of magnetic interchange reconnection at the boundaries of the coronal hole. In the stabilization phase, the magnetic properties and area of the coronal hole became constant. The EM of the coronal hole decreased, which we interpret as a reduction in plasma density due to the onset of plasma outflow into interplanetary space. As the coronal hole rotated towards the solar limb, it merged with a nearby pre-existing coronal hole. At the next solar rotation, the coronal hole was still apparent, indicating a lifetime of >1 solar rotation.

astro-ph.SR↗

Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers

Chain-of-thought reasoning and scratchpads have emerged as critical tools for enhancing the computational capabilities of transformers. While theoretical results show that polynomial-length scratchpads can extend transformers' expressivity from $TC^0$ to $PTIME$, their required length remains poorly understood. Empirical evidence even suggests that transformers need scratchpads even for many problems in $TC^0$, such as Parity or Multiplication, challenging optimistic bounds derived from circuit complexity. In this work, we initiate the study of systematic lower bounds for the number of chain-of-thought steps across different algorithmic problems, in the hard-attention regime. We study a variety of algorithmic problems, and provide bounds that are tight up to logarithmic factors. Overall, these results contribute to emerging understanding of the power and limitations of chain-of-thought reasoning.

cs.LG↗

One Size Fits None: Rethinking Fairness in Medical AI

Machine learning (ML) models are increasingly used to support clinical decision-making. However, real-world medical datasets are often noisy, incomplete, and imbalanced, leading to performance disparities across patient subgroups. These differences raise fairness concerns, particularly when they reinforce existing disadvantages for marginalized groups. In this work, we analyze several medical prediction tasks and demonstrate how model performance varies with patient characteristics. While ML models may demonstrate good overall performance, we argue that subgroup-level evaluation is essential before integrating them into clinical workflows. By conducting a performance analysis at the subgroup level, differences can be clearly identified-allowing, on the one hand, for performance disparities to be considered in clinical practice, and on the other hand, for these insights to inform the responsible development of more effective models. Thereby, our work contributes to a practical discussion around the subgroup-sensitive development and deployment of medical ML models and the interconnectedness of fairness and transparency.

cs.LG↗

Position: Pause Recycling LoRAs and Prioritize Mechanisms to Uncover Limits and Effectiveness

Merging or routing low-rank adapters (LoRAs) has emerged as a popular solution for enhancing large language models, particularly when data access is restricted by regulatory or domain-specific constraints. This position paper argues that the research community should shift its focus from developing new merging or routing algorithms to understanding the conditions under which reusing LoRAs is truly effective. Through theoretical analysis and synthetic two-hop reasoning and math word-problem tasks, we examine whether reusing LoRAs enables genuine compositional generalization or merely reflects shallow pattern matching. Evaluating two data-agnostic methods--parameter averaging and dynamic adapter selection--we found that reusing LoRAs often fails to logically integrate knowledge across disjoint fine-tuning datasets, especially when such knowledge is underrepresented during pretraining. Our empirical results, supported by theoretical insights into LoRA's limited expressiveness, highlight the preconditions and constraints of reusing them for unseen tasks and cast doubt on its feasibility as a truly data-free approach. We advocate for pausing the pursuit of novel methods for recycling LoRAs and emphasize the need for rigorous mechanisms to guide future academic research in adapter-based model merging and practical system designs for practitioners.

cs.CL↗

A Formal Framework for Understanding Length Generalization in Transformers

A major challenge for transformers is generalizing to sequences longer than those observed during training. While previous works have empirically shown that transformers can either succeed or fail at length generalization depending on the task, theoretical understanding of this phenomenon remains limited. In this work, we introduce a rigorous theoretical framework to analyze length generalization in causal transformers with learnable absolute positional encodings. In particular, we characterize those functions that are identifiable in the limit from sufficiently long inputs with absolute positional encodings under an idealized inference scheme using a norm-based regularizer. This enables us to prove the possibility of length generalization for a rich family of problems. We experimentally validate the theory as a predictor of success and failure of length generalization across a range of algorithmic and formal language tasks. Our theory not only explains a broad set of empirical observations but also opens the way to provably predicting length generalization capabilities in transformers.

cs.LG↗

Velocity and Density Fluctuations in the Quiet Sun Corona

We investigate the properties and relationship between Doppler-velocity fluctuations and intensity fluctuations in the off-limb quiet Sun corona. These are expected to reflect the properties of Alfvenic and compressive waves, respectively. The data come from the Coronal Multichannel Polarimeter (COMP). These data were studied using spectral methods to estimate the power spectra, amplitudes, perpendicular correlation lengths, phases, trajectories, dispersion relations, and propagation speeds of both types of fluctuations. We find that most velocity fluctuations are due to Alfvenic waves, but that intensity fluctuations come from a variety of sources, likely including fast and slow mode waves, as well as aperiodic variations. The relation between the velocity and intensity fluctuations differs depending on the underlying coronal structure. On short closed loops, the velocity and intensity fluctuations have similar power spectra and speeds. In contrast, on longer nearly radial trajectories, the velocity and intensity fluctuations have different power spectra, with the velocity fluctuations propagating at much faster speeds than the intensity fluctuations. Considering the temperature sensitivity of COMP, these longer structures are more likely to be closed fields lines of the quiet Sun rather than cooler open field lines. That is, we find the character of the interactions of Alfvenic waves and density fluctuations depends on the length of the magnetic loop on which they are traveling.

astro-ph.SR↗

Elemental Abundances at Coronal Hole Boundaries as a Means to Investigate Interchange Reconnection and the Solar Wind

The origin of the slow solar wind is not well understood, unlike the fast solar wind which originates from coronal holes. In-situ elemental abundances of the slow solar wind suggest that it originates from initially closed field lines that become open. Coronal hole boundary regions are a potential source of slow solar wind as there open field lines interact with the closed loops through interchange reconnection. Our primary aim is to quantify the role of interchange reconnection at the boundaries of coronal holes. To this end, we have measured the relative abundances of different elements at these boundaries. Reconnection is expected to modulate the relative abundances through the first ionization potential (FIP) effect. For our analysis we used spectroscopic data from the extreme ultraviolet imaging spectrometer (EIS) on board Hinode. To account for the temperature structure of the observed region we computed the differential emission measure (DEM). Using the DEM we were able to infer the ratio between coronal and photospheric abundances, known as the FIP bias. By examining the variation of the FIP bias moving from the coronal hole to the quiet Sun, we have been able to constrain models of interchange reconnection. The FIP bias variation in the boundary region around the coronal hole has an approximate width of 30-50 Mm, comparable to the size of supergranules. This boundary region is also a source of open flux into interplanetary space. We find that there is an additional ~30% open flux that originates from this boundary region.

astro-ph.SR↗

Revised Point-Spread Functions for the Atmospheric Imaging Assembly onboard the Solar Dynamics Observatory

We present revised point-spread functions (PSFs) for the Atmospheric Imaging Assembly (AIA) onboard the Solar Dynamics Observatory (SDO). These PSFs provide a robust estimate of the light diffracted by the meshes holding the entrance and focal plane filters and the light that is diffusely scattered over medium- to long-distance by the micro-roughness of the mirrors. We first calibrate the diffracted light using flare images. Our modeling of the diffracted light provides reliable determinations of the mesh parameters and finds that about 24 to 33% of the collected light is diffracted, depending on the AIA channel. Then, we fit for the diffuse scattered light using partially lunar occulted images. We find that the diffuse scattered light can be modeled as a superposition of two power law functions that scatter light over the entire length of the detector. The amount of diffuse scattered light ranges from 10 to 35 %, depending on the AIA channel. In total, AIA diffracts and scatters about 40 to 60 % of the collected light over medium to long distances. When correcting for this, bright image regions increase in intensity by about 30 %, dark image regions decrease in intensity by up to 90 %, and the associated differential emission measure analysis of solar features are affected accordingly. Finally, we compare the image reconstructions using our new PSFs to those using the AIA team PSFs and the PSFs of Poduval et al. (2013). We find that our PSFs outperform the others; our PSFs correct well for the flare diffraction pattern and predict accurately the long-distance scattered light in lunar occultations.

astro-ph.SR↗

Measurement of energy reduction by inertial Alfvén waves propagating through parallel gradients in the Alfvén speed

We have studied the propagation of inertial Alfvén waves through parallel gradients in the Alfvén speed using the Large Plasma Device at the University of California, Los Angeles. The reflection and transmission of Alfvén waves through inhomogeneities in the background plasma is important for understanding wave propagation, turbulence, and heating in space, laboratory, and astrophysical plasmas. Here we \rev{present inertial Alfvén waves, under conditions relevant to solar flares and the solar corona. We find} that the transmission of the inertial Alfvén waves is reduced as the sharpness of the gradient is increased. Any reflected waves were below the detection limit of our experiment and reflection cannot account for all of the energy not transmitted through the gradient. Our findings indicate that, for both kinetic and inertial Alfvén waves, the controlling parameter for the transmission of the waves through an Alfvén speed gradient is the ratio of the Alfvén wavelength along the gradient divided by the scale length of the gradient. Furthermore, our results suggest that an as-yet-unidentified damping process occurs in the gradient.

physics.plasm-ph↗

Emergent Stack Representations in Modeling Counter Languages Using Transformers

Transformer architectures are the backbone of most modern language models, but understanding the inner workings of these models still largely remains an open problem. One way that research in the past has tackled this problem is by isolating the learning capabilities of these architectures by training them over well-understood classes of formal languages. We extend this literature by analyzing models trained over counter languages, which can be modeled using counter variables. We train transformer models on 4 counter languages, and equivalently formulate these languages using stacks, whose depths can be understood as the counter values. We then probe their internal representations for stack depths at each input token to show that these models when trained as next token predictors learn stack-like representations. This brings us closer to understanding the algorithmic details of how transformers learn languages and helps in circuit discovery.

cs.CL↗

InversionView: A General-Purpose Method for Reading Information from Neural Activations

The inner workings of neural networks can be better understood if we can fully decipher the information encoded in neural activations. In this paper, we argue that this information is embodied by the subset of inputs that give rise to similar activations. We propose InversionView, which allows us to practically inspect this subset by sampling from a trained decoder model conditioned on activations. This helps uncover the information content of activation vectors, and facilitates understanding of the algorithms implemented by transformer models. We present four case studies where we investigate models ranging from small transformers to GPT-2. In these studies, we show that InversionView can reveal clear information contained in activations, including basic information about tokens appearing in the context, as well as more complex information, such as the count of certain tokens, their relative positions, and abstract knowledge about the subject. We also provide causally verified circuits to confirm the decoded information.

cs.LG↗

High-Resolution Laboratory Measurements of M-shell Fe EUV Line Emission using EBIT-I

Solar physicists routinely utilize observations of Ar-like Fe IX and Cl-like Fe X emission to study a variety of solar structures. However, unidentified lines exist in the Fe IX and Fe X spectra, greatly impeding the spectroscopic diagnostic potential of these ions. Here, we present measurements using the Lawrence Livermore National Laboratory EBIT-I electron beam ion trap in the wavelength range 238-258 A. These studies enable us to unambiguously identify the charge state associated with each of the observed lines. This wavelength range is of particular interest because it contains the Fe IX density diagnostic line ratio 241.74 A/244.91 A, which is predicted to be one of the best density diagnostics of the solar corona, as well as the Fe X 257.26 A magnetic-field-induced transition. We compare our measurements to the Fe IX and Fe X lines tabulated in CHIANTI v10.0.1, which is used for modeling the solar spectrum. In addition, we have measured previously unidentified Fe X lines that will need to be added to CHIANTI and other spectroscopic databases.

astro-ph.SR↗

Separations in the Representational Capabilities of Transformers and Recurrent Architectures

Transformer architectures have been widely adopted in foundation models. Due to their high inference costs, there is renewed interest in exploring the potential of efficient recurrent architectures (RNNs). In this paper, we analyze the differences in the representational capabilities of Transformers and RNNs across several tasks of practical relevance, including index lookup, nearest neighbor, recognizing bounded Dyck languages, and string equality. For the tasks considered, our results show separations based on the size of the model required for different architectures. For example, we show that a one-layer Transformer of logarithmic width can perform index lookup, whereas an RNN requires a hidden state of linear size. Conversely, while constant-size RNNs can recognize bounded Dyck languages, we show that one-layer Transformers require a linear size for this task. Furthermore, we show that two-layer Transformers of logarithmic size can perform decision tasks such as string equality or disjointness, whereas both one-layer Transformers and recurrent models require linear size for these tasks. We also show that a log-size two-layer Transformer can implement the nearest neighbor algorithm in its forward pass; on the other hand recurrent models require linear size. Our constructions are based on the existence of $N$ nearly orthogonal vectors in $O(\log N)$ dimensional space and our lower bounds are based on reductions from communication complexity problems. We supplement our theoretical results with experiments that highlight the differences in the performance of these architectures on practical-size sequences.

cs.LG↗