SearcharxivSearch

arXiv subjects

Robin Karlsson

Publications and source records attributed to Robin Karlsson.

At least 19 recordsLinked to original sources

OPE = QNM

At microscopic scales the linear response of thermal states of large-$N$ CFTs is governed by a thermal operator product expansion (OPE), while at large scales response is governed by collective excitations known as quasinormal modes (QNM). We show that the OPE and QNM representations of the retarded correlator in mixed time and spatial momentum coordinates have an overlapping region of convergence in the complex time plane, giving a map between OPE and QNM data. We show that large-overtone QNM asymptotics are related to OPE singularities, while low-overtone QNM data appear in analytic continuation from short to large times. Using this approach we obtain new analytic results for QNM asymptotics, and numerically obtain low-overtone QNMs from OPE data for the Schwarzschild-AdS$_5$ black brane. We further show that the OPE spectrum is intimately related to QNM data through a set of sum rules which we derive in Mellin space. Finally, using the lightcone OPE, we argue that stress tensor correlators at large spatial momentum thermalise slower as the conformal collider bounds approach saturation. This work points to a new thermal bootstrap programme where OPE and QNM data constrain each other.

hep-th

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations

Deploying massive large language models (LLMs) as continuous cognitive engines for robotics is bottlenecked by the time-to-first-token (TTFT) latency required to process extensive state histories. Existing solutions like RAG or sliding windows compromise global context or incur prohibitive re-computation costs. We formalize the optimal task structure for minimizing latency and theoretically prove that prefix stability, incremental extensibility, and asynchronous state reconciliation are necessary conditions for real-time performance. Building on these proofs, we introduce the Cached State Representation (CSR) framework as the practical instantiation of these properties, ensuring optimal KV-cache reuse. To sustain these properties over infinite horizons, we further propose an Asynchronous State Reconciliation (ASR) algorithm that offloads state memory eviction to a parallel computational resource to eliminate latency spikes. On a physical robot wirelessly connected to an on-premise GPU server, CSR achieves a 26-fold latency reduction (14.67s to 0.56s) for 120K token contexts with a 235B parameter model compared to a standard baseline. On an embodied AI benchmark, we achieve SOTA recall (0.836 vs. 0.459) while maintaining RAG-level latency. ASR is validated to sustain bounded, spike-free TTFT over 10 eviction cycles in continuous real-world operation. Together, CSR and ASR enable massive LLMs to function as continuously operating, high-frequency (> 2 Hz) embodied policies.

cs.RO

Conformal collider bootstrap in ${\mathcal N}=4$ SYM

We use a combination of perturbation theory, holography, supersymmetric localization, integrability, and numerical conformal bootstrap methods to constrain the energy-energy correlator in $\text{SU}(N_c)$ ${\mathcal N}=4$ SYM at finite coupling. For finite $N_c$, we derive lower bounds on the second and fourth multipoles of the energy-energy correlator at different couplings, along with a smeared energy-energy correlator as a function of the angle between the two detectors. We present evidence that our lower bounds on the multipoles are nearly saturated by the ${\cal N} = 4$ SYM theory. In the planar limit, we further use dispersive functionals to obtain tight two-sided bounds on both the first three non-trivial multipoles and on the angular dependence of the energy-energy correlator. As the coupling is varied from weak to strong, the energy-energy correlator exhibits a transition from single-trace to double-trace operator dominance in the collinear limit, which we characterize quantitatively. A similar phenomenon occurs in QCD, where a parton-hadron transition is observed as detectors are brought closer together.

hep-th

Bouncing off a stringy singularity

A sharp signature of the black hole singularity in holography is a divergence in the boundary thermal two-point function at a specific point in the complex time plane. This divergence arises from a null geodesic that bounces off the black hole singularity. At finite 't Hooft coupling, stringy corrections to the bulk dynamics cannot be neglected, and the fate of the bouncing geodesic is an open question. We propose a simple scenario in which the singularity in the two-point function is shifted slightly into the complex plane, thereby smoothing it out into a finite-size bump. We demonstrate this smoothing explicitly in a microscopic example, namely the Sachdev-Ye-Kitaev model at infinite temperature, where the correlator is under analytic control. Our result suggests a bulk description of planar theories at finite coupling as stringy black holes.

hep-th

MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction

Pedestrian action prediction is of great significance for many applications such as autonomous driving. However, state-of-the-art methods lack explainability to make trustworthy predictions. In this paper, a novel framework called MulCPred is proposed that explains its predictions based on multi-modal concepts represented by training samples. Previous concept-based methods have limitations including: 1) they cannot directly apply to multi-modal cases; 2) they lack locality to attend to details in the inputs; 3) they suffer from mode collapse. These limitations are tackled accordingly through the following approaches: 1) a linear aggregator to integrate the activation results of the concepts into predictions, which associates concepts of different modalities and provides ante-hoc explanations of the relevance between the concepts and the predictions; 2) a channel-wise recalibration module that attends to local spatiotemporal regions, which enables the concepts with locality; 3) a feature regularization loss that encourages the concepts to learn diverse patterns. MulCPred is evaluated on multiple datasets and tasks. Both qualitative and quantitative results demonstrate that MulCPred is promising in improving the explainability of pedestrian action prediction without obvious performance degradation. Furthermore, by removing unrecognizable concepts from MulCPred, the cross-dataset prediction performance is improved, indicating the feasibility of further generalizability of MulCPred.

cs.CV

Energy correlations and Planckian collisions

Energy correlations characterize the energy flux through detectors at infinity produced in a collision event. Remarkably, in holographic conformal field theories, they probe high-energy gravitational scattering in the dual anti-de Sitter geometry. We use known properties of high-energy gravitational scattering and its unitarization to explore the leading quantum-gravity correction to the energy-energy correlator at strong coupling. We find that it includes a part originating from large impact parameter scattering that is non-analytic in the angle between detectors and is $\log N_c$ enhanced compared to the standard $1/N_c$ perturbative expansion. It is sensitive to the full bulk geometry, including the internal manifold, providing a refined probe of the emergent holographic spacetime. Similarly, scattering at small impact parameters leads to contributions that are further enhanced by extra powers of the 't Hooft coupling assuming it is corrected by stringy effects. We conclude that energy correlations are sensitive to the UV properties of the dual gravitational theory and thus provide a promising target for the conformal bootstrap.

hep-th

Black hole bulk-cone singularities

Lorentzian correlators of local operators exhibit surprising singularities in theories with gravity duals. These are associated with null geodesics in an emergent bulk geometry. We analyze singularities of the thermal response function dual to propagation of waves on the AdS Schwarzschild black hole background. We derive the analytic form of the leading singularity dual to a bulk geodesic that winds around the black hole. Remarkably, it exhibits a boundary group velocity larger than the speed of light, whose dual is the angular velocity of null geodesics at the photon sphere. The strength of this singularity is controlled by the classical Lyapunov exponent associated with the instability of nearly bound photon orbits. In this sense, the bulk-cone singularity can be identified as the universal feature that encodes the ubiquitous black hole photon sphere in a dual holographic CFT. To perform the computation analytically, we express the two-point correlator as an infinite sum over Regge poles, and then evaluate this sum using WKB methods. We also compute the smeared correlator numerically, which in particular allows us to check and support our analytic predictions. We comment on the resolution of black hole bulk-cone singularities by stringy and gravitational effects into black hole bulk-cone "bumps". We conclude that these bumps are robust, and could serve as a target for simulations of black hole-like geometries in table-top experiments.

hep-th

Compositional Semantics for Open Vocabulary Spatio-semantic Representations

Vision-language models (VLMs) transform environment percepts into vision-language semantics interpretable by LLMs. However, completing complex tasks often requires reasoning about information beyond what is currently perceived. We propose latent compositional semantic embeddings z* as a principled learning-based knowledge representation for queryable spatio-semantic memories. We mathematically prove that z* can always be found, and that the optimal z* is the centroid for any set Z. We derive a probabilistic bound for estimating separability of related and unrelated semantics. We prove that z* is discoverable from visual appearance and singular descriptions by iterative gradient descent. We experimentally verify our findings on four embedding spaces including CLIP and SBERT. Our results show that z* can represent up to 10 semantics encoded by SBERT, and up to 100 semantics for ideal uniformly distributed high-dimensional embeddings. We introduce three new datasets with overlapping semantics to show that common VLMs trained on conventional nonoverlapping annotations discover z*. Our novel sufficient similarity inference method overcomes fundamental limitations of conventional inference, and improves higher-level overlapping semantic inference performance by 19.63 mIoU on average.

cs.CV

R-Cut: Enhancing Explainability in Vision Transformers with Relationship Weighted Out and Cut

Transformer-based models have gained popularity in the field of natural language processing (NLP) and are extensively utilized in computer vision tasks and multi-modal models such as GPT4. This paper presents a novel method to enhance the explainability of Transformer-based image classification models. Our method aims to improve trust in classification results and empower users to gain a deeper understanding of the model for downstream tasks by providing visualizations of class-specific maps. We introduce two modules: the ``Relationship Weighted Out" and the ``Cut" modules. The ``Relationship Weighted Out" module focuses on extracting class-specific information from intermediate layers, enabling us to highlight relevant features. Additionally, the ``Cut" module performs fine-grained feature decomposition, taking into account factors such as position, texture, and color. By integrating these modules, we generate dense class-specific visual explainability maps. We validate our method with extensive qualitative and quantitative experiments on the ImageNet dataset. Furthermore, we conduct a large number of experiments on the LRN dataset, specifically designed for automatic driving danger alerts, to evaluate the explainability of our method in complex backgrounds. The results demonstrate a significant improvement over previous methods. Moreover, we conduct ablation experiments to validate the effectiveness of each module. Through these experiments, we are able to confirm the respective contributions of each module, thus solidifying the overall effectiveness of our proposed approach.

cs.CV

Thermal Stress Tensor Correlators near Lightcone and Holography

We consider thermal stress-tensor two-point functions in holographic theories in the near-lightcone regime and analyse them using the operator product expansion (OPE). In the limit we consider only the leading-twist multi-stress tensors contribute and the correlators depend on a particular combination of lightcone momenta. We argue that such correlators are described by three universal functions, which can be holographically computed in Einstein gravity; higher-derivative terms in the gravitational Lagrangian enter the arguments of these functions via the cubic stress-tensor couplings and the thermal stress-tensor expectation value in the dual CFT. We compute the retarded correlators and observe that in addition to the perturbative OPE, which contributes to the real part, there is a non-perturbative contribution to the imaginary part.

hep-th

Learning to Predict Navigational Patterns from Partial Observations

Human beings cooperatively navigate rule-constrained environments by adhering to mutually known navigational patterns, which may be represented as directional pathways or road lanes. Inferring these navigational patterns from incompletely observed environments is required for intelligent mobile robots operating in unmapped locations. However, algorithmically defining these navigational patterns is nontrivial. This paper presents the first self-supervised learning (SSL) method for learning to infer navigational patterns in real-world environments from partial observations only. We explain how geometric data augmentation, predictive world modeling, and an information-theoretic regularizer enables our model to predict an unbiased local directional soft lane probability (DSLP) field in the limit of infinite data. We demonstrate how to infer global navigational patterns by fitting a maximum likelihood graph to the DSLP field. Experiments show that our SSL model outperforms two SOTA supervised lane graph prediction models on the nuScenes dataset. We propose our SSL method as a scalable and interpretable continual learning paradigm for navigation by perception. Code is available at https://github.com/robin-karlsson0/dslp.

cs.CV

A thermal product formula

We show that holographic thermal two-sided two-point correlators take the form of a product over quasi-normal modes (QNMs). Due to this fact, the two-point function admits a natural dispersive representation with a positive discontinuity at the location of QNMs. We explore the general constraints on the structure of QNMs that follow from the operator product expansion, the presence of the singularity inside the black hole, and the hydrodynamic expansion of the correlator. We illustrate these constraints through concrete examples. We suggest that the product formula for thermal correlators may hold for more general large N chaotic systems, and we check this hypothesis in several models.

hep-th

Predictive World Models from Real-World Partial Observations

Cognitive scientists believe adaptable intelligent agents like humans perform reasoning through learned causal mental simulations of agents and environments. The problem of learning such simulations is called predictive world modeling. Recently, reinforcement learning (RL) agents leveraging world models have achieved SOTA performance in game environments. However, understanding how to apply the world modeling approach in complex real-world environments relevant to mobile robots remains an open question. In this paper, we present a framework for learning a probabilistic predictive world model for real-world road environments. We implement the model using a hierarchical VAE (HVAE) capable of predicting a diverse set of fully observed plausible worlds from accumulated sensor observations. While prior HVAE methods require complete states as ground truth for learning, we present a novel sequential training method to allow HVAEs to learn to predict complete states from partially observed states only. We experimentally demonstrate accurate spatial structure prediction of deterministic regions achieving 96.21 IoU, and close the gap to perfect prediction by 62% for stochastic regions using the best prediction. By extending HVAEs to cases where complete ground truth states do not exist, we facilitate continual learning of spatial prediction as a step towards realizing explainable and comprehensive predictive world models for real-world mobile robotics applications. Code is available at https://github.com/robin-karlsson0/predictive-world-models.

cs.CV

Freedom near Lightcone and ANEC Saturation

Averaged Null Energy Conditions (ANECs) hold in unitary quantum field theories. In conformal field theories, ANECs in states created by the application of the stress tensor to the vacuum lead to three constraints on the stress-tensor three-point couplings, depending on the choice of polarization. The same constraints follow from considering two-point functions of the stress tensor in a thermal state and focusing on the contribution of the stress tensor in the operator product expansion (OPE). One can observe this in holographic Gauss-Bonnet gravity, where ANEC saturation coincides with the appearance of superluminal signal propagation in thermal states. We show that, when this happens, the corresponding generalizations of ANECs for higher-spin multi-stress tensor operators with minimal twist are saturated as well and all contributions from such operators to the thermal two-point functions vanish in the lightcone limit. This leads to a special near-lightcone behavior of the thermal stress-tensor correlators -- they take the vacuum form, independent of temperature.

hep-th

Thermal Stress Tensor Correlators, OPE and Holography

In strongly coupled conformal field theories with a large central charge important light degrees of freedom are the stress tensor and its composites, multi-stress tensors. We consider the OPE expansion of two-point functions of the stress tensor in thermal and heavy states and focus on the contributions from the stress tensor and double-stress tensors in four spacetime dimensions. We compare the results to the holographic finite temperature two-point functions and read off conformal data beyond the leading order in the large central charge expansion. In particular, we compute corrections to the OPE coefficients which determine the near-lightcone behavior of the correlators. We also compute the anomalous dimensions of the double-stress tensor operators.

hep-th

ViCE: Improving Dense Representation Learning by Superpixelization and Contrasting Cluster Assignment

Recent self-supervised models have demonstrated equal or better performance than supervised methods, opening for AI systems to learn visual representations from practically unlimited data. However, these methods are typically classification-based and thus ineffective for learning high-resolution feature maps that preserve precise spatial information. This work introduces superpixels to improve self-supervised learning of dense semantically rich visual concept embeddings. Decomposing images into a small set of visually coherent regions reduces the computational complexity by $\mathcal{O}(1000)$ while preserving detail. We experimentally show that contrasting over regions improves the effectiveness of contrastive learning methods, extends their applicability to high-resolution images, improves overclustering performance, superpixels are better than grids, and regional masking improves performance. The expressiveness of our dense embeddings is demonstrated by improving the SOTA unsupervised semantic segmentation benchmark on Cityscapes, and for convolutional models on COCO.

cs.CV

CFT correlators, ${\cal W}$-algebras and Generalized Catalan Numbers

In two spacetime dimensions the Virasoro heavy-heavy-light-light (HHLL) vacuum block in a certain limit is governed by the Catalan numbers. The equation for their generating function can be generalized to a differential equation which the logarithm of the block satisfies. We show that a similar story holds for the HHLL ${\cal W}_N$ vacuum blocks, where a suitable generalization of the Catalan numbers plays the main role. Moreover, the ${\cal W}_N$ blocks have the same form as the stress tensor sector of HHLL near lightcone conformal correlators in $2(N-1)$ spacetime dimensions. In the latter case the Catalan numbers are generalized to the numbers of linear extensions of certain partially ordered sets.

hep-th

Learning a Model for Inferring a Spatial Road Lane Network Graph using Self-Supervision

Interconnected road lanes are a central concept for navigating urban roads. Currently, most autonomous vehicles rely on preconstructed lane maps as designing an algorithmic model is difficult. However, the generation and maintenance of such maps is costly and hinders large-scale adoption of autonomous vehicle technology. This paper presents the first self-supervised learning method to train a model to infer a spatially grounded lane-level road network graph based on a dense segmented representation of the road scene generated from onboard sensors. A formal road lane network model is presented and proves that any structured road scene can be represented by a directed acyclic graph of at most depth three while retaining the notion of intersection regions, and that this is the most compressed representation. The formal model is implemented by a hybrid neural and search-based model, utilizing a novel barrier function loss formulation for robust learning from partial labels. Experiments are conducted for all common road intersection layouts. Results show that the model can generalize to new road layouts, unlike previous approaches, demonstrating its potential for real-world application as a practical learning-based lane-level map generator.

cs.CV