SearcharxivSearch

arXiv subjects

Raul Gomez

Publications and source records attributed to Raul Gomez.

17 recordsLinked to original sources

A Stone-von Neumann equivalence of categories for smooth representations of the Heisenberg group

The classical Stone-von Neuman theorem relates the irreducible unitary representations of the Heisenberg group $H_n$ to non-trivial unitary characters of its center $Z$, and plays a crucial role in the construction of the oscillator representation for the metaplectic group. In this paper we extend these ideas to non-unitary and non-irreducible representations, thereby obtaining an equivalence of categories between certain representations of $Z$ and those of $H_n$. Our main result is a smooth equivalence, which involves the fundamental ideas of du Cloux on differentiable representations and smooth imprimitivity systems for Nash groups. We show how to extend the oscillator representation to the smooth setting and give an application to degenerate Whittaker models for representations of reductive groups. We also include an algebraic equivalence, which can be regarded as a generalization of Kashiwara's lemma from the theory of $D$-modules.

math.RT

Retrieval Guided Unsupervised Multi-domain Image-to-Image Translation

Image to image translation aims to learn a mapping that transforms an image from one visual domain to another. Recent works assume that images descriptors can be disentangled into a domain-invariant content representation and a domain-specific style representation. Thus, translation models seek to preserve the content of source images while changing the style to a target visual domain. However, synthesizing new images is extremely challenging especially in multi-domain translations, as the network has to compose content and style to generate reliable and diverse images in multiple domains. In this paper we propose the use of an image retrieval system to assist the image-to-image translation task. First, we train an image-to-image translation model to map images to multiple domains. Then, we train an image retrieval model using real and generated images to find images similar to a query one in content but in a different domain. Finally, we exploit the image retrieval system to fine-tune the image-to-image translation model and generate higher quality images. Our experiments show the effectiveness of the proposed solution and highlight the contribution of the retrieval network, which can benefit from additional unlabeled data and help image-to-image translation models in the presence of scarce data.

cs.CV

Whittaker supports for representations of reductive groups

Let $F$ be either $\mathbb{R}$ or a finite extension of $\mathbb{Q}_p$, and let $G$ be a finite central extension of the group of $F$-points of a reductive group defined over $F$. Also let $π$ be a smooth representation of $G$ (Frechet of moderate growth if $F=\mathbb{R}$). For each nilpotent orbit $\mathcal{O}$ we consider a certain Whittaker quotient $π_{\mathcal{O}}$ of $π$. We define the Whittaker support WS$(π)$ to be the set of maximal $\mathcal{O}$ among those for which $π_{\mathcal{O}}\neq 0$. In this paper we prove that all $\mathcal{O}\in\mathrm{WS}(π)$ are quasi-admissible nilpotent orbits, generalizing some of the results in [Moe96,JLS16]. If $F$ is $p$-adic and $π$ is quasi-cuspidal then we show that all $\mathcal{O}\in\mathrm{WS}(π)$ are $F$-distinguished, i.e. do not intersect the Lie algebra of any proper Levi subgroup of $G$ defined over $F$. We also give an adaptation of our argument to automorphic representations, generalizing some results from [GRS03,Shen16,JLS16,Cai] and confirming some conjectures from [Ginz06]. Our methods are a synergy of the methods of the above-mentioned papers, and of our preceding paper [GGS17].

math.RT

Location Sensitive Image Retrieval and Tagging

People from different parts of the globe describe objects and concepts in distinct manners. Visual appearance can thus vary across different geographic locations, which makes location a relevant contextual information when analysing visual data. In this work, we address the task of image retrieval related to a given tag conditioned on a certain location on Earth. We present LocSens, a model that learns to rank triplets of images, tags and coordinates by plausibility, and two training strategies to balance the location influence in the final ranking. LocSens learns to fuse textual and location information of multimodal queries to retrieve related images at different levels of location granularity, and successfully utilizes location information to improve image tagging.

cs.CV

Exploring Hate Speech Detection in Multimodal Publications

In this work we target the problem of hate speech detection in multimodal publications formed by a text and an image. We gather and annotate a large scale dataset from Twitter, MMHS150K, and propose different models that jointly analyze textual and visual information for hate speech detection, comparing them with unimodal detection. We provide quantitative and qualitative results and analyze the challenges of the proposed task. We find that, even though images are useful for the hate speech detection task, current multimodal models cannot outperform models analyzing only text. We discuss why and open the field and the dataset for further research.

cs.CV

Selective Style Transfer for Text

This paper explores the possibilities of image style transfer applied to text maintaining the original transcriptions. Results on different text domains (scene text, machine printed text and handwritten text) and cross modal results demonstrate that this is feasible, and open different research lines. Furthermore, two architectures for selective style transfer, which means transferring style to only desired image pixels, are proposed. Finally, scene text selective style transfer is evaluated as a data augmentation technique to expand scene text detection datasets, resulting in a boost of text detectors performance. Our implementation of the described models is publicly available.

cs.CV

Self-Supervised Learning from Web Data for Multimodal Retrieval

Self-Supervised learning from multimodal image and text data allows deep neural networks to learn powerful features with no need of human annotated data. Web and Social Media platforms provide a virtually unlimited amount of this multimodal data. In this work we propose to exploit this free available data to learn a multimodal image and text embedding, aiming to leverage the semantic knowledge learnt in the text domain and transfer it to a visual model for semantic image retrieval. We demonstrate that the proposed pipeline can learn from images with associated textwithout supervision and analyze the semantic structure of the learnt joint image and text embedding space. We perform a thorough analysis and performance comparison of five different state of the art text embeddings in three different benchmarks. We show that the embeddings learnt with Web and Social Media data have competitive performances over supervised methods in the text based image retrieval task, and we clearly outperform state of the art in the MIRFlickr dataset when training in the target data. Further, we demonstrate how semantic multimodal image retrieval can be performed using the learnt embeddings, going beyond classical instance-level retrieval problems. Finally, we present a new dataset, InstaCities1M, composed by Instagram images and their associated texts that can be used for fair comparison of image-text embeddings.

cs.CV

Learning to Learn from Web Data through Deep Semantic Embeddings

In this paper we propose to learn a multimodal image and text embedding from Web and Social Media data, aiming to leverage the semantic knowledge learnt in the text domain and transfer it to a visual model for semantic image retrieval. We demonstrate that the pipeline can learn from images with associated text without supervision and perform a thourough analysis of five different text embeddings in three different benchmarks. We show that the embeddings learnt with Web and Social Media data have competitive performances over supervised methods in the text based image retrieval task, and we clearly outperform state of the art in the MIRFlickr dataset when training in the target data. Further we demonstrate how semantic multimodal image retrieval can be performed using the learnt embeddings, going beyond classical instance-level retrieval problems. Finally, we present a new dataset, InstaCities1M, composed by Instagram images and their associated texts that can be used for fair comparison of image-text embeddings.

cs.CV

Learning from #Barcelona Instagram data what Locals and Tourists post about its Neighbourhoods

Massive tourism is becoming a big problem for some cities, such as Barcelona, due to its concentration in some neighborhoods. In this work we gather Instagram data related to Barcelona consisting on images-captions pairs and, using the text as a supervisory signal, we learn relations between images, words and neighborhoods. Our goal is to learn which visual elements appear in photos when people is posting about each neighborhood. We perform a language separate treatment of the data and show that it can be extrapolated to a tourists and locals separate analysis, and that tourism is reflected in Social Media at a neighborhood level. The presented pipeline allows analyzing the differences between the images that tourists and locals associate to the different neighborhoods. The proposed method, which can be extended to other cities or subjects, proves that Instagram data can be used to train multi-modal (image and text) machine learning models that are useful to analyze publications about a city at a neighborhood level. We publish the collected dataset, InstaBarcelona and the code used in the analysis.

cs.CV

TextTopicNet - Self-Supervised Learning of Visual Features Through Embedding Images on Semantic Text Spaces

The immense success of deep learning based methods in computer vision heavily relies on large scale training datasets. These richly annotated datasets help the network learn discriminative visual features. Collecting and annotating such datasets requires a tremendous amount of human effort and annotations are limited to popular set of classes. As an alternative, learning visual features by designing auxiliary tasks which make use of freely available self-supervision has become increasingly popular in the computer vision community. In this paper, we put forward an idea to take advantage of multi-modal context to provide self-supervision for the training of computer vision algorithms. We show that adequate visual features can be learned efficiently by training a CNN to predict the semantic textual context in which a particular image is more probable to appear as an illustration. More specifically we use popular text embedding techniques to provide the self-supervision for the training of deep CNN. Our experiments demonstrate state-of-the-art performance in image classification, object detection, and multi-modal retrieval compared to recent self-supervised or naturally-supervised approaches.

cs.CV

Improving Text Proposals for Scene Images with Fully Convolutional Networks

Text Proposals have emerged as a class-dependent version of object proposals - efficient approaches to reduce the search space of possible text object locations in an image. Combined with strong word classifiers, text proposals currently yield top state of the art results in end-to-end scene text recognition. In this paper we propose an improvement over the original Text Proposals algorithm of Gomez and Karatzas (2016), combining it with Fully Convolutional Networks to improve the ranking of proposals. Results on the ICDAR RRC and the COCO-text datasets show superior performance over current state-of-the-art.

cs.CV

Generalized and degenerate Whittaker models

We study generalized and degenerate Whittaker models for reductive groups over local fields of characteristic zero (archimedean or non-archimedean). Our main result is the construction of epimorphisms from the generalized Whittaker model corresponding to a nilpotent orbit to any degenerate Whittaker model corresponding to the same orbit, and to certain degenerate Whittaker models corresponding to bigger orbits. We also give choice-free definitions of generalized and degenerate Whittaker models. Finally, we explain how our methods imply analogous results for Whittaker-Fourier coefficients of automorphic representations. For $\mathrm{GL}_n(F)$ this implies that a smooth admissible representation $π$ has a generalized Whittaker model $\mathcal{W}_{\mathcal{O}}(π)$ corresponding to a nilpotent coadjoint orbit $\mathcal{O}$ if and only if $\mathcal{O}$ lies in the (closure of) the wave-front set $\mathrm{WF}(π)$. Previously this was only known to hold for $F$ non-archimedean and $\mathcal{O}$ maximal in $\mathrm{WF}(π)$, see [MW87]. We also express $\mathcal{W}_{\mathcal{O}}(π)$ as an iteration of a version of the Bernstein-Zelevinsky derivatives [BZ77,AGS15a]. This enables us to extend to $\mathrm{GL_n}(\mathbb{R})$ and $\mathrm{GL_n}(\mathbb{C})$ several further results from [MW87] on the dimension of $\mathcal{W}_{\mathcal{O}}(π)$ and on the exactness of the generalized Whittaker functor.

math.RT

Local theta lifting of generalized Whittaker models associated to nilpotent orbits

Let $(G,\tilde{G})$ be a reductive dual pair over a local field ${\Fontauri k}$ of characteristic 0, and denote by $V$ and $\tilde{V}$ the standard modules of $G$ and $\tilde{G}$, respectively. Consider the set $Max Hom(V,\tilde{V})$ of full rank elements in $Hom(V,\tilde{V})$, and the nilpotent orbit correspondence $\mathcal{O} \subset \mathfrak{g}$ and $Θ(\mathcal{O})\subset \tilde{\mathfrak{g}}$ induced by elements of $Max Hom(V,\tilde{V})$ via the moment maps. Let $(π,\mathscr{V})$ be a smooth irreducible representation of $G$. We show that there is a correspondence of the generalized Whittaker models of $π$ of type $\mathcal{O}$ and of $Θ(π)$ of type $Θ(\mathcal{O})$, where $Θ(π)$ is the full theta lift of $π$. When $(G,\tilde{G})$ is in the stable range with $G$ the smaller member, every nilpotent orbit $\mathcal{O} \subset \mathfrak{g}$ is in the image of the moment map from $Max Hom (V,\tilde{V})$. In this case, and for ${\Fontauri k}$ non-Archimedean, the result has been previously obtained by Mœglin in a different approach.

math.RT

The Bessel-Plancherel theorem and applications

Let $G$ be a simple Lie Group with finite center, and let $K\subset G$ be a maximal compact subgroup. We say that $G$ is a Lie group of tube type if $G/K$ is a hermitian symmetric space of tube type. For such a Lie group $G$, we can find a parabolic subgroup $P=MAN$, with given Langlands decomposition, such that $N$ is abelian, and $N$ admits a generic character with compact stabilizer. We will call any parabolic subgroup $P$ satisfying this properties a Siegel parabolic. Let $(π,V)$ be an admissible, smooth, Fréchet representation of a Lie group of tube type $G$, and let $P \subset G$ be a Siegel parabolic subgroup. If $χ$ is a generic character of $N$, let $Wh_χ(V)={λ:V \longrightarrow \mathbb{C} | λ(π(n)v)=χ(n)v}$ be the space of Bessel models of $V$. After describing the classification of all the simple Lie groups of tube type, we will give a characterization of the space of Bessel models of an induced representation. As a corollary of this characterization we obtain a local multiplicity one theorem for the space of Bessel models of an irreducible representation of $G$. As an application of this results we calculate the Bessel-Plancherel measure of a Lie group of tube type, $L^2(N\backslash G;χ)$, where $χ$ is a generic character of $N$. Then we use Howe's theory of dual pairs to show that the Plancherel measure of the space $L^2(O(p-r,q-s)\backslash O(p,q))$ is the pullback, under the $Θ$ lift, of the Bessel-Plancherel measure $L^2(N\backslash Sp(m,\mathbb{R});χ)$, where $m=r+s$ and $χ$ is a generic character that depends on $r$ and $s$.

math.RT

A Conjecture of Sakellaridis-Venkatesh on the Unitary Spectrum of Spherical Varieties

In a recent preprint, Sakellaridis and Venkatesh considered the spectral decomposition of the space $L^2(X)$, where $X = H\G$ is a spherical variety and $G$ is a real or $p$-adic group, and stated a conjecture describing this decomposition in terms of a dual group $\check{G}_X$ associated to $X$. The main purpose of this paper is to verify the above conjecture in many cases when $X$, has low rank. In particular, we demonstrate this conjecture for many cases when $X$ has rank 1, and also some cases when $X$ has rank 2 or 3.

math.RT

Bessel Models for General Admissible Induced Representations: The Compact Stabilizer Case

A holomorphic continuation of Jacquet type integrals for parabolic subgroups with abelian nilradical is studied. Complete results are given for generic characters with compact stabilizer and arbitrary representations induced from admissible representations. A description of all of the pertinent examples is given. These results give a complete description of the Bessel models corresponding to compact stabilizer.

math.RT

Determinant Expansions of Signed Matrices and of Certain Jacobians

This paper treats two topics: matrices with sign patterns and Jacobians of certain mappings. The main topic is counting the number of plus and minus coefficients in the determinant expansion of sign patterns and of these Jacobians. The paper is motivated by an approach to chemical networks initiated by Craciun and Feinberg. We also give a graph-theoretic test for determining when the Jacobian of a chemical reaction dynamics has a sign pattern.

math.RA