Searcharxiv⌕ Search

arXiv subjects

René Wagner

Publications and source records attributed to René Wagner.

6 recordsLinked to original sources

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch representations with text embeddings of corresponding concepts. In this work, we investigate this critical issue and propose novel techniques to enhance this capability in foundational vision-language models. First, we reveal that a patch-level distillation procedure significantly boosts dense patch-text alignment -- surprisingly, the patch-text alignment of the distilled student model strongly surpasses that of the teacher model. This observation inspires us to consider modifications to pretraining recipes, leading us to propose iBOT++, an upgrade to the commonly-used iBOT masked image objective, where unmasked tokens also contribute directly to the loss. This dramatically enhances patch-text alignment of pretrained models. Additionally, to improve vision-language pretraining efficiency and effectiveness, we modify the exponential moving average setup in the learning recipe, and introduce a caption sampling strategy to benefit from synthetic captions at different granularities. Combining these components, we develop TIPSv2, a new family of image-text encoder models suitable for a wide range of downstream applications. Through comprehensive experiments on 9 tasks and 20 datasets, we demonstrate strong performance, generally on par with or better than recent vision encoder models. Code and models are released via our project page at https://gdm-tipsv2.github.io/ .

cs.CV↗

Dichroic Electron Emission Patterns from Oriented Helium Ions

We report a joint experimental and theoretical study using a combination of polarization-controlled free-electron-laser (FEL) and near-infra\-red (NIR) pulses in a synchronized two-color photo\-ionization scheme. Excited He$^+$ ions, created by extreme ultraviolet (XUV) circularly polarized radiation from the XUV-FEL FERMI in the oriented $3p\, (m\!=\!+1)$ state, are exposed to circularly polarized 784-nm NIR radiation with peak intensities from $10^{12}\,\rm W/cm^2$ to $\rm 10^{13}\,W/cm^2$. The angular distribution of the ejected electrons exhibit a strong dichroism depending on the NIR intensity. While the co-rotating case is defined by a single path, for the counter-rotating case, there are two dominant pathways whose relative strength and phase difference are determined.

physics.atom-ph↗

Wavelength-Dependent Photodissociation of Iodomethylbutane

Ultrashort XUV pulses of the Free-Electron-LASer in Hamburg (FLASH) were used to investigate laser-induced fragmentation patterns of the prototypical chiral molecule 1-iodo-2-methyl-butane (C$_5$H$_{11}$I) in a pump-probe scheme. Ion velocity-map images and mass spectra of optical-laser-induced fragmentation were obtained for subsequent FEL exposure with photon energies of 63 eV and 75 eV. These energies specifically address the iodine 4d edge of neutral and singly charged iodine, respectively. The presented ion spectra for two optical pump-laser wavelengths, i.e., 800 nm and 267 nm, reveal substantially different cationic fragment yields in dependence on the wavelength and intensity. For the case of 800-nm-initiated fragmentation, the molecule dissociates notably slower than for the 267-nm pump. The results underscore the importance of considering optical-laser wavelength and intensity in the dissociation dynamics of this prototypical chiral molecule that is a promising candidate for future studies of its asymmetric nature.

physics.chem-ph↗

Learning the RoPEs: Better 2D and 3D Position Encodings with STRING

We introduce STRING: Separable Translationally Invariant Position Encodings. STRING extends Rotary Position Encodings, a recently proposed and widely used algorithm in large language models, via a unifying theoretical framework. Importantly, STRING still provides exact translation invariance, including token coordinates of arbitrary dimensionality, whilst maintaining a low computational footprint. These properties are especially important in robotics, where efficient 3D token representation is key. We integrate STRING into Vision Transformers with RGB(-D) inputs (color plus optional depth), showing substantial gains, e.g. in open-vocabulary object detection and for robotics controllers. We complement our experiments with a rigorous mathematical analysis, proving the universality of our methods.

cs.LG↗

Linear Transformer Topological Masking with Graph Random Features

When training transformers on graph-structured data, incorporating information about the underlying topology is crucial for good performance. Topological masking, a type of relative position encoding, achieves this by upweighting or downweighting attention depending on the relationship between the query and keys in a graph. In this paper, we propose to parameterise topological masks as a learnable function of a weighted adjacency matrix -- a novel, flexible approach which incorporates a strong structural inductive bias. By approximating this mask with graph random features (for which we prove the first known concentration bounds), we show how this can be made fully compatible with linear attention, preserving $\mathcal{O}(N)$ time and space complexity with respect to the number of input tokens. The fastest previous alternative was $\mathcal{O}(N \log N)$ and only suitable for specific graphs. Our efficient masking algorithms provide strong performance gains for tasks on image and point cloud data, including with $>30$k nodes.

cs.LG↗

Integrating Generic Sensor Fusion Algorithms with Sound State Representations through Encapsulation of Manifolds

Common estimation algorithms, such as least squares estimation or the Kalman filter, operate on a state in a state space S that is represented as a real-valued vector. However, for many quantities, most notably orientations in 3D, S is not a vector space, but a so-called manifold, i.e. it behaves like a vector space locally but has a more complex global topological structure. For integrating these quantities, several ad-hoc approaches have been proposed. Here, we present a principled solution to this problem where the structure of the manifold S is encapsulated by two operators, state displacement [+]:S x R^n --> S and its inverse [-]: S x S --> R^n. These operators provide a local vector-space view δ; --> x [+] δ; around a given state x. Generic estimation algorithms can then work on the manifold S mainly by replacing +/- with [+]/[-] where appropriate. We analyze these operators axiomatically, and demonstrate their use in least-squares estimation and the Unscented Kalman Filter. Moreover, we exploit the idea of encapsulation from a software engineering perspective in the Manifold Toolkit, where the [+]/[-] operators mediate between a "flat-vector" view for the generic algorithm and a "named-members" view for the problem specific functions.

cs.RO↗