SearcharxivSearch

arXiv subjects

Ashish Thapliyal

Publications and source records attributed to Ashish Thapliyal.

3 recordsLinked to original sources

PaLI: A Jointly-Scaled Multilingual Language-Image Model

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of language and vision. PaLI generates text based on visual and textual inputs, and with this interface performs many vision, language, and multimodal tasks, in many languages. To train PaLI, we make use of large pre-trained encoder-decoder language models and Vision Transformers (ViTs). This allows us to capitalize on their existing capabilities and leverage the substantial cost of training them. We find that joint scaling of the vision and language components is important. Since existing Transformers for language are much larger than their vision counterparts, we train a large, 4-billion parameter ViT (ViT-e) to quantify the benefits from even larger-capacity vision models. To train PaLI, we create a large multilingual mix of pretraining tasks, based on a new image-text training set containing 10B images and texts in over 100 languages. PaLI achieves state-of-the-art in multiple vision and language tasks (such as captioning, visual question-answering, scene-text understanding), while retaining a simple, modular, and scalable design.

cs.CV

Denoising Large-Scale Image Captioning from Alt-text Data using Content Selection Models

Training large-scale image captioning (IC) models demands access to a rich and diverse set of training examples, gathered from the wild, often from noisy alt-text data. However, recent modeling approaches to IC often fall short in terms of performance in this case, because they assume a clean annotated dataset (as opposed to the noisier alt-text--based annotations), and employ an end-to-end generation approach, which often lacks both controllability and interpretability. We address these problems by breaking down the task into two simpler, more controllable tasks -- skeleton prediction and skeleton-based caption generation. Specifically, we show that selecting content words as skeletons} helps in generating improved and denoised captions when leveraging rich yet noisy alt-text--based uncurated datasets. We also show that the predicted English skeletons can be further cross-lingually leveraged to generate non-English captions, and present experimental results covering caption generation in French, Italian, German, Spanish and Hindi. We also show that skeleton-based prediction allows for better control of certain caption properties, such as length, content, and gender expression, providing a handle to perform human-in-the-loop semi-automatic corrections.

cs.CL

Entanglement of Assistance

The newfound importance of ``entanglement as a resource'' in quantum computation and quantum communication compels us to quantify it in as many distinct ways as possible. Here we explore a new measure of entanglement for mixed quantum states of bipartite systems, which we name the Entanglement of Assistance. We show it to be the maximum average entanglement of all pure-state ensembles consistent with the given density matrix. In this sense, the Entanglement of Assistance is a quantity directly dual to the more standard Entanglement of Formation. With the help of lower and upper bounds, we calculate the Entanglement of Assistance for a few cases and use these results to show that it possesses the surprising property of superadditivity. We believe that this may shed some light on the question of additivity for the Entanglement of Formation.

quant-ph