SearcharxivSearch

arXiv subjects

Thomas Klein

Publications and source records attributed to Thomas Klein.

15 recordsLinked to original sources

Low-Pass Filtering Improves Behavioral Alignment of Vision Models

Despite their impressive performance on computer vision benchmarks, Deep Neural Networks (DNNs) still fall short of adequately modeling human visual behavior, as measured by error consistency and shape bias. Recent work hypothesized that behavioral alignment can be drastically improved through \emph{generative} -- rather than \emph{discriminative} -- classifiers, with far-reaching implications for models of human vision. Here, we instead show that the increased alignment of generative models can be largely explained by a seemingly innocuous resizing operation in the generative model which effectively acts as a low-pass filter. In a series of controlled experiments, we show that removing high-frequency spatial information from discriminative models like CLIP drastically increases their behavioral alignment. Simply blurring images at test-time -- rather than training on blurred images -- achieves a new state-of-the-art score on the model-vs-human benchmark, halving the current alignment gap between DNNs and human observers. Furthermore, low-pass filters are likely optimal, which we demonstrate by directly optimizing filters for alignment. To contextualize the performance of optimal filters, we compute the frontier of all possible pareto-optimal solutions to the benchmark, which was formerly unknown. We explain our findings by observing that the frequency spectrum of optimal Gaussian filters roughly matches the spectrum of band-pass filters implemented by the human visual system. We show that the contrast sensitivity function, describing the inverse of the contrast threshold required for humans to detect a sinusoidal grating as a function of spatiotemporal frequency, is approximated well by Gaussian filters of the specific width that also maximizes error consistency.

cs.CV

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation. This shift has sparked interest in using intermediate visualizations as a reasoning aid, akin to human mental imagery. Central to this idea is the ability to form, maintain, and manipulate visual representations in a goal-oriented manner. To evaluate and probe this capability, we develop MentisOculi, a procedural, stratified suite of multi-step reasoning problems amenable to visual solution, tuned to challenge frontier models. Evaluating visual strategies ranging from latent tokens to explicit generated imagery, we find they generally fail to improve performance. Analysis of UMMs specifically exposes a critical limitation: While they possess the textual reasoning capacity to solve a task and can sometimes generate correct visuals, they suffer from compounding generation errors and fail to leverage even ground-truth visualizations. Our findings suggest that despite their inherent appeal, visual thoughts do not yet benefit model reasoning. MentisOculi establishes the necessary foundation to analyze and close this gap across diverse model families.

cs.AI

Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers

Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and human observers is error consistency (EC). EC allows for more fine-grained comparisons of behavior than other metrics such as accuracy, and has been used in the influential Brain-Score benchmark to rank different DNNs by their behavioral consistency with humans. Previously, EC values have been reported without confidence intervals. However, empirically measured EC values are typically noisy -- thus, without confidence intervals, valid benchmarking conclusions are problematic. Here we improve on standard EC in two ways: First, we show how to obtain confidence intervals for EC using a bootstrapping technique, allowing us to derive significance tests for EC. Second, we propose a new computational model relating the EC between two classifiers to the implicit probability that one of them copies responses from the other. This view of EC allows us to give practical guidance to scientists regarding the number of trials required for sufficiently powerful, conclusive experiments. Finally, we use our methodology to revisit popular NeuroAI-results. We find that while the general trend of behavioral differences between humans and machines holds up to scrutiny, many reported differences between deep vision models are statistically insignificant. Our methodology enables researchers to design adequately powered experiments that can reliably detect behavioral differences between models, providing a foundation for more rigorous benchmarking of behavioral alignment.

q-bio.NC

LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

Out-of-distribution (OOD) robustness is a desired property of computer vision models. Improving model robustness requires high-quality signals from robustness benchmarks to quantify progress. While various benchmark datasets such as ImageNet-C were proposed in the ImageNet era, most ImageNet-C corruption types are no longer OOD relative to today's large, web-scraped datasets, which already contain common corruptions such as blur or JPEG compression artifacts. Consequently, these benchmarks are no longer well-suited for evaluating OOD robustness in the era of web-scale datasets. Indeed, recent models show saturating scores on ImageNet-era OOD benchmarks, indicating that it is unclear whether models trained on web-scale datasets truly become better at OOD generalization or whether they have simply been exposed to the test distortions during training. To address this, we introduce LAION-C as a benchmark alternative for ImageNet-C. LAION-C consists of six novel distortion types specifically designed to be OOD, even for web-scale datasets such as LAION. In a comprehensive evaluation of state-of-the-art models, we find that the LAION-C dataset poses significant challenges to contemporary models, including MLLMs such as Gemini and GPT-4o. We additionally conducted a psychophysical experiment to evaluate the difficulty of our corruptions for human observers, enabling a comparison of models to lab-quality human robustness data. We observe a paradigm shift in OOD generalization: from humans outperforming models, to the best models now matching or outperforming the best human observers.

cs.CV

Ab initio modeling of TWIP and TRIP effects in $\beta$-Ti alloys

Transformations in bcc-$\beta$, hcp-$\alpha$, and the $\omega$ phases of Ti alloys are studied using Density Functional Theory for pure Ti and Ti alloyed with Al, Si, V, Cr, Fe, Cu, Nb, Mo, and Sn. The $\beta$-stabilization caused by alloying Si, Fe, Cr, and Mo was observed, but the most stable phase appears between the $\beta$ and the $\alpha$ phases, corresponding to the martensitic $\alpha''$ phase. Next, the $\{112\}\langle11\bar1\rangle$ bcc twins are separated by a positive barrier, which further increases by alloying w.r.t. pure Ti. The $\{332\}\langle11\bar3\rangle$ twinning yields negative barriers for all species but Mo and Fe. This is because the transition state is structurally similar to the $\alpha$ phase, which is preferred over the $\beta$ phase for the majority of alloying elements. Lastly, the impact of alloying on twin boundary energies is discussed. These results may serve as design guidelines for novel Ti-based alloys with specific application areas.

cond-mat.mtrl-sci

How Aligned are Different Alignment Metrics?

In recent years, various methods and benchmarks have been proposed to empirically evaluate the alignment of artificial neural networks to human neural and behavioral data. But how aligned are different alignment metrics? To answer this question, we analyze visual data from Brain-Score (Schrimpf et al., 2018), including metrics from the model-vs-human toolbox (Geirhos et al., 2021), together with human feature alignment (Linsley et al., 2018; Fel et al., 2022) and human similarity judgements (Muttenthaler et al., 2022). We find that pairwise correlations between neural scores and behavioral scores are quite low and sometimes even negative. For instance, the average correlation between those 80 models on Brain-Score that were fully evaluated on all 69 alignment metrics we considered is only 0.198. Assuming that all of the employed metrics are sound, this implies that alignment with human perception may best be thought of as a multidimensional concept, with different methods measuring fundamentally different aspects. Our results underline the importance of integrative benchmarking, but also raise questions about how to correctly combine and aggregate individual metrics. Aggregating by taking the arithmetic average, as done in Brain-Score, leads to the overall performance currently being dominated by behavior (95.25% explained variance) while the neural predictivity plays a less important role (only 33.33% explained variance). As a first step towards making sure that different alignment metrics all contribute fairly towards an integrative benchmark score, we therefore conclude by comparing three different aggregation options.

q-bio.NC

Scale Alone Does not Improve Mechanistic Interpretability in Vision Models

In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic interpretability for a diverse suite of nine models and find no scaling effect for interpretability - neither for model nor dataset size. Specifically, none of the investigated state-of-the-art models are easier to interpret than the GoogLeNet model from almost a decade ago. Latest-generation vision models appear even less interpretable than older architectures, hinting at a regression rather than improvement, with modern models sacrificing interpretability for accuracy. These results highlight the need for models explicitly designed to be mechanistically interpretable and the need for more helpful interpretability methods to increase our understanding of networks at an atomic level. We release a dataset containing more than 130'000 human responses from our psychophysical evaluation of 767 units across nine models. This dataset facilitates research on automated instead of human-based interpretability evaluations, which can ultimately be leveraged to directly optimize the mechanistic interpretability of models.

cs.CV

Analysis of FDML lasers with meter range coherence

FDML lasers provide sweep rates in the MHz range at wide optical bandwidths, making them ideal sources for high speed OCT. Recently, at lower speed, ultralong-range swept-source OCT has been demonstrated1, 2 using a tunable vertical cavity surface emitting laser (VCSEL) and also using a Vernier-tunable laser. These sources provide relatively high sweep rates and meter range coherence lengths. In order to achieve similar coherence, we developed an extremely well dispersion compensated Fourier Domain Mode Locked (FDML) laser, running at 3.2 MHz sweep rate and 120 nm spectral bandwidth. We demonstrate that this laser offers meter range coherence and enables volumetric long range OCT of moving objects.

physics.ins-det

Shot-Noise limited Time-encoded (TICO) Raman spectroscopy

Raman scattering, an inelastic scattering mechanism, provides information about molecular excitation energies and can be used to identify chemical compounds. Albeit being a powerful analysis tool, especially for label-free biomedical imaging with molecular contrast, it suffers from inherently low signal levels. This practical limitation can be overcome by non-linear enhancement techniques like stimulated Raman scattering (SRS). In SRS, an additional light source stimulates the Raman scattering process. This can lead to orders of magnitude increase in signal levels and hence faster acquisition in biomedical imaging. However, achieving a broad spectral coverage in SRS is technically challenging and the signal is no longer background-free, as either stimulated Raman gain (SRG) or loss (SRL) is measured, turning a sensitivity limit into a dynamic range limit. Thus, the signal has to be isolated from the laser background light, requiring elaborate methods for minimizing detection noise. Here we analyze the detection sensitivity of a shot-noise limited broadband stimulated time-encoded Raman (TICO-Raman) system in detail. In time-encoded Raman, a wavelength-swept Fourier Domain Mode Locked (FDML) laser covers a broad range of Raman transition energies while allowing a dual-balanced detection for lowering the detection noise to the fundamental shot-noise limit.

physics.optics

Megahertz FDML Laser with up to 143nm Sweep Range for Ultrahigh Resolution OCT at 1050nm

We present a new design of a Fourier Domain Mode Locked laser (FDML laser), which provides a new record in sweep range at ~1um center wavelength: At the fundamental sweep rate of 2x417 kHz we reach 143nm bandwidth and 120nm with 4x buffering at 1.67MHz sweep rate. The latter configuration of our system is characterized: The FWHM of the point spread function (PSF) of a mirror is 5.6um (in tissue). Human in vivo retinal imaging is performed with the MHz laser showing more details in vascular structures. Here we could measure an axial resolution of 6.0um by determining the FWHM of specular reflex in the image. Additionally, challenges related to such a high sweep bandwidth such as water absorption are investigated.

physics.optics

Flexible A-scan rate MHz OCT: Computational downscaling by coherent averaging

In order to realize fast OCT-systems with adjustable line rate, we investigate averaging of image data from an FDML based MHz-OCT-system. The line rate can be reduced in software and traded in for increased system sensitivity and image quality. We compare coherent and incoherent averaging to effectively scale down the system speed of a 3.2 MHz FDML OCT system to around 100 kHz in postprocessing. We demonstrate that coherent averaging is possible with MHz systems without special interferometer designs or digital phase stabilisation. We show OCT images of a human finger knuckle joint in vivo with very high quality and deep penetration.

physics.med-ph

Preferential site occupancy of alloying elements in TiAl-based phases

First principles calculations are used to study the preferential occupation of ternary alloying additions into the binary Ti-Al phases, namely $\gamma$-TiAl, $\alpha_2$-Ti$_3$Al, $\beta_{\mathrm{o}}$-TiAl, and B19-TiAl. While the early transition metals (TMs, group IVB, VB , and VIB elements) prefer to substitute for Ti atoms in the $\gamma$-, $\alpha_2$-, and B19-phases, they preferentially occupy Al sites in the $\beta_{\mathrm{o}}$-TiAl. Si is in this context an anomaly, as it prefers to sit on the Al sublattice for all four phases. B and C are shown to prefer octahedral Ti-rich interstitial positions instead of substitutional incorporation. The site preference energy is linked with the alloying-induced changes of energy of formation, hence alloying-related (de)stabilisation of the phases. We further show that the phase-stabilisation effect of early TMs on $\beta_{\mathrm{o}}$-phase has a different origin depending on their valency. Finally, an extensive comparison of our predictions with available theoretical and experimental data (which is, however, limited mostly to the $\gamma$-phase) shows a consistent picture.

cond-mat.mtrl-sci

First supra-THz Heterodyne Array Receivers for Astronomy with the SOFIA Observatory

We present the upGREAT THz heterodyne arrays for far-infrared astronomy. The Low Frequency Array (LFA) is designed to cover the 1.9-2.5 THz range using 2x7-pixel waveguide-based HEB mixer arrays in a dual polarization configuration. The High Frequency Array (HFA) will perform observations of the [OI] line at ~4.745 THz using a 7-pixel waveguide-based HEB mixer array. This paper describes the common design for both arrays, cooled to 4.5 K using closed- cycle pulse tube technology. We then show the laboratory and telescope characterization of the first array with its 14 pixels (LFA), which culminated in the successful commissioning in May 2015 aboard the SOFIA airborne observatory observing the [CII] fine structure transition at 1.905 THz. This is the first successful demonstration of astronomical observations with a heterodyne focal plane array above 1 THz and is also the first time high- power closed-cycle coolers for temperatures below 4.5 K are operated on an airborne platform.

astro-ph.IM

Time-Encoded Raman: Fiber-based, hyperspectral, broadband stimulated Raman microscopy

Raman sensing and Raman microscopy are amongst the most specific optical technologies to identify the chemical compounds of unknown samples, and to enable label-free biomedical imaging with molecular contrast. However, the high cost and complexity, low speed, and incomplete spectral information provided by current technology are major challenges preventing more widespread application of Raman systems. To overcome these limitations, we developed a new method for stimulated Raman spectroscopy and Raman imaging using continuous wave (CW), rapidly wavelength swept lasers. Our all-fiber, time-encoded Raman (TICO-Raman) setup uses a Fourier Domain Mode Locked (FDML) laser source to achieve a unique combination of high speed, broad spectral coverage (750 cm-1 - 3150 cm-1) and high resolution (0.5 cm-1). The Raman information is directly encoded and acquired in time. We demonstrate quantitative chemical analysis of a solvent mixture and hyperspectral Raman microscopy with molecular contrast of plant cells.

physics.optics

Dynamics and PDR properties in IC1396A

We investigate the gas dynamics and the physical properties of photodissociation regions (PDRs) in IC1396A, which is an illuminated bright-rimmed globule with internal structures created by young stellar objects. Our mapping observations of the [CII] emission in IC1396A with GREAT onboard SOFIA revealed the detailed velocity structure of this region. We combined them with observations of the [CI] 3P_1 - 3P_0 and CO(4-3) emissions to study the dynamics of the different tracers and physical properties of the PDRs. The [CII] emission generally matches the IRAC 8 micron, which traces the polycyclic aromatic hydrocarbon (PAH) emissions. The CO(4-3) emission peaks inside the globule, and the [CI] emission is strong in outer regions, following the 8 micron emission to some degree, but its peak is different from that of [CII]. The [CII] emitting gas shows a clear velocity gradient within the globule, which is not significant in the [CI] and CO(4-3) emission. Some clumps that are prominent in [CII] emission appear to be blown away from the rim of the globule. The observed ratios of [CII]/[CI] and [CII]/CO(4-3) are compared to the KOSMA-tau PDR model, which indicates a density of 10^4-10^5 cm-3.

astro-ph.GA