Searcharxiv⌕ Search

arXiv subjects

Yu Shi

Publications and source records attributed to Yu Shi.

At least 145 records · Page 8Linked to original sources

High dimensional AdS-like black hole and Phase transition in Einstein-bumblebee gravity

In this paper we obtain an exact high dimensional anti-de Sitter (AdS) black hole solution in Einstein-bumblebee gravity theory. This AdS-like black hole can only exist with a linear functional potential of the bumblebee field. We find that the Smarr formula and the first law of black hole thermodynamics can still be constructed in this Lorentz symmetry breaking black hole spacetime as long as its temperature, entropy and volume are slightly modified. We find also that there exist two kinds of phase transition: small-large black hole phase transition and Hawking-Page phase transition, like those of Schwarzschild AdS black hole. After Lorentz symmetry breaking, the black hole mass at divergent point of heat capacity becomes small, and the Gibbs free energy of the meta-stable large black hole is also smaller, showing that the large stable black hole can be more easily formed.

gr-qc↗

Quantum Fourier transform on photonic qubits using cavity QED

We propose a quantum Fourier transform on photons in which a single atom-coupled cavity system mediates the photon-photon interactions. Our protocol utilizes time-delay feedback of photons and requires no active feedforward control. The time-delay feedback enables a single atom-cavity system to implement a quantum Fourier transform on an arbitrary number of photonic qubits on-the-fly, while rapid tuning of the atomic transition implements arbitrary controlled-phase gates. We analyze the performance of the protocol numerically and show that it can implement quantum Fourier transforms with tens of photons using state-of-the-art cavity quantum electrodynamics.

quant-ph↗

Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering

Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore the visual appearance-motion information at different temporal scales, it is unknown how to incorporate the multilevel processing capacity of a deep learning model with such multiscale information. Targeting these issues, this paper proposes a novel Multilevel Hierarchical Network (MHN) with multiscale sampling for VideoQA. MHN comprises two modules, namely Recurrent Multimodal Interaction (RMI) and Parallel Visual Reasoning (PVR). With a multiscale sampling, RMI iterates the interaction of appearance-motion information at each scale and the question embeddings to build the multilevel question-guided visual representations. Thereon, with a shared transformer encoder, PVR infers the visual cues at each level in parallel to fit with answering different question types that may rely on the visual information at relevant levels. Through extensive experiments on three VideoQA datasets, we demonstrate improved performances than previous state-of-the-arts and justify the effectiveness of each part of our method.

cs.CV↗

i-Code: An Integrative and Composable Multimodal Learning Framework

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel attention mechanisms and other architectural innovations to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five video understanding tasks and the GLUE NLP benchmark, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.

cs.LG↗

On the kinematic constraint in BFKL near threshold region

We introduce a modified Balitskii-Fadin-Kuraev-Lipatov (BFKL) equation with the rapidity veto that originates from external kinematic constraint. Though it is a formally sub-leading power contribution, such kinematic effect becomes very important especially in the threshold region, where there is no sufficient phase space for the small $x$ evolution to be fully developed. We also investigate the phenomenological consequences of the kinematic constraint BFKL equation in the forward particle production processes in pp and eA collisions. It is shown that particle yield at large transverse momentum is significantly suppressed in the threshold region.

hep-ph↗

An Empirical Study of Graphormer on Large-Scale Molecular Modeling Datasets

This technical note describes the recent updates of Graphormer, including architecture design modifications, and the adaption to 3D molecular dynamics simulation. The "Graphormer-V2" could attain better results on large-scale molecular modeling datasets than the vanilla one, and the performance gain could be consistently obtained on downstream tasks. In addition, we show that with a global receptive field and an adaptive aggregation strategy, Graphormer is more powerful than classic message-passing-based GNNs. Graphormer-V2 achieves much less MAE than the vanilla Graphormer on the PCQM4M quantum chemistry dataset used in KDD Cup 2021, where the latter one won the first place in this competition. In the meanwhile, Graphormer-V2 greatly outperforms the competitors in the recent Open Catalyst Challenge, which is a competition track on NeurIPS 2021 workshop, and aims to model the catalyst-adsorbate reaction system with advanced AI models. All models could be found at \url{https://github.com/Microsoft/Graphormer}.

physics.chem-ph↗

Digital quantum simulation and Pseudoquantum Simulation of $\mathbb{Z}_2$ Gauge Higgs Model

We present a quantum algorithm for digital quantum simulation of the $\mathbb{Z}_2$ gauge-Higgs model on a $3\times 3$ lattice, which is based on Trotter decomposition, the quantum adiabatic algorithm and its circuit realization. Then we perform a classical demonstration, dubbed a pseudoquantum simulation, on a GPU simulator. We obtain useful results on this model, which suggest the topological properties of the deconfined phase and help to clarify the phase diagram. It is suggested that the tricitical point, where the second-order critical lines of deconfinement-confinement transition and of deconfinement-Higgs transition meet, seems to be on the the first-order critical line of confinement-Higgs transition, at a point other than the end of this critical line.

hep-lat↗

Solving Hamiltonian Cycle Problem using Quantum $\mathbb{Z}_2$ Lattice Gauge Theory

The Hamiltonian cycle (HC) problem in graph theory is a well-known NP-complete problem. We present an approach in terms of $\mathbb{Z}_2$ lattice gauge theory (LGT) defined on the lattice with the graph as its dual. When the coupling parameter $g$ is less than the critical value $g_c$, the ground state is a superposition of all configurations with closed strings of spins in a same single-spin state, which can be obtained by using an adiabatic quantum algorithm with time complexity $O(\frac{1}{g_c^2} \sqrt{ \frac{1}{\varepsilon} N_e^{3/2}(N_v^3 + \frac{N_e}{g_c}}))$, where $N_v$ and $N_e$ are the numbers of vertices and edges of the graph respectively. A subsequent search for a HC among those closed-strings solves the HC problem. For some random samples of small graphs, we demonstrate that the dependence of the average value of $g_c$ on $\sqrt{N_{hc}}$, $N_{hc}$ being the number of HCs, and that of the average value of $\frac{1}{g_c}$ on $N_e$ are both linear. It is thus suggested that for some graphs, the HC problem may be solved in polynomial time. A possible quantum algorithm using $g_c$ to infer $N_{hc}$ is also discussed.

quant-ph↗

Higgs mode in $2+1$ dimensional $O(2)$ model

We investigate the spectral functions of the Higgs mode in $O(2)$ model, which can be experimentally realized in a 2D Bose gas. Zero temperature limit is considered. Our calculation fully includes the 2-loop contributions. Peaks show up in the spectral functions of both the longitudinal and the scalar susceptibilities. Thus this model cannot explain the disappearance of the response at the weak interaction limit. Neither it can explain the similarity between the longitudinal and the scalar susceptibilities in the visibility of the Higgs mode. A possible lower peak at about $2m_σ$ is also noted.

hep-ph↗

Pursuing the Precision Study for Color Glass Condensate in Forward Hadron Productions

With the tremendous accomplishments of RHIC and the LHC experiments and the advent of the future Electron-Ion Collider on the horizon, the quest for compelling evidence of the color glass condensate (CGC) has become one of the most aspiring goals in the high energy Quantum Chromodynamics research. Pursuing this question requires developing the precision test of the CGC formalism. By systematically implementing the threshold resummation, we significantly improve the stability of the next-to-leading-order calculation in CGC for forward rapidity hadron productions in $pp$ and $pA$ collisions, especially in the high $p_T$ region, and obtain reliable descriptions of all existing data measured at RHIC and the LHC across all $p_T$ regions. Consequently, this technique can pave the way for the precision studies of the CGC next-to-leading-order predictions by confronting them with a large amount of precise data.

hep-ph↗

Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition

The sparsely-gated Mixture of Experts (MoE) can magnify a network capacity with a little computational complexity. In this work, we investigate how multi-lingual Automatic Speech Recognition (ASR) networks can be scaled up with a simple routing algorithm in order to achieve better accuracy. More specifically, we apply the sparsely-gated MoE technique to two types of networks: Sequence-to-Sequence Transformer (S2S-T) and Transformer Transducer (T-T). We demonstrate through a set of ASR experiments on multiple language data that the MoE networks can reduce the relative word error rates by 16.3% and 4.6% with the S2S-T and T-T, respectively. Moreover, we thoroughly investigate the effect of the MoE on the T-T architecture in various conditions: streaming mode, non-streaming mode, the use of language ID and the label decoder with the MoE.

cs.CL↗

Effective field theory of the Higgs mode in a two-dimensional dilute Bose gas

We investigate the spectral function of the Higgs mode in a two dimensional Bose gas, by using the effective field theory in the zero temperature limit. Our approach explains the experimental feature that the peak of the spectral function is a soft continuum rather than a sharp peak, broadened and vanishing in the superfluid phase, which cannot be explained in terms of the $O(2)$ model. We also find that the scalar susceptibility is the same as the longitudinal susceptibility.

cond-mat.quant-gas↗

Florence: A New Foundation Model for Computer Vision

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications. While existing vision foundation models such as CLIP, ALIGN, and Wu Dao 2.0 focus mainly on mapping images and textual representations to a cross-modal shared representation, we introduce a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine (object), from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth). By incorporating universal visual-language representations from Web-scale image-text data, our Florence model can be easily adapted for various computer vision tasks, such as classification, retrieval, object detection, VQA, image caption, video retrieval and action recognition. Moreover, Florence demonstrates outstanding performance in many types of transfer learning: fully sampled fine-tuning, linear probing, few-shot transfer and zero-shot transfer for novel images and objects. All of these properties are critical for our vision foundation model to serve general purpose vision tasks. Florence achieves new state-of-the-art results in majority of 44 representative benchmarks, e.g., ImageNet-1K zero-shot classification with top-1 accuracy of 83.74 and the top-5 accuracy of 97.18, 62.4 mAP on COCO fine tuning, 80.36 on VQA, and 87.8 on Kinetics-600.

cs.CV↗

Optimizing Alignment of Speech and Language Latent Spaces for End-to-End Speech Recognition and Understanding

The advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the performance of AED-based E2E ASR is to introduce an extra text encoder for leveraging extensive text data and thus capture more context-aware linguistic information. However, this approach brings a mismatch problem between the speech encoder and the text encoder due to the different units used for modeling. In this paper, we propose an embedding aligner and modality switch training to better align the speech and text latent spaces. The embedding aligner is a shared linear projection between text encoder and speech encoder trained by masked language modeling (MLM) loss and connectionist temporal classification (CTC), respectively. The modality switch training randomly swaps speech and text embeddings based on the forced alignment result to learn a joint representation space. Experimental results show that our proposed approach achieves a relative 14% to 19% word error rate (WER) reduction on Librispeech ASR task. We further verify its effectiveness on spoken language understanding (SLU), i.e., an absolute 2.5% to 2.8% F1 score improvement on SNIPS slot filling task.

cs.SD↗

Temporal Pyramid Transformer with Multimodal Interaction for Video Question Answering

Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language understanding. While existing approaches seldom leverage the appearance-motion information in the video at multiple temporal scales, the interaction between the question and the visual information for textual semantics extraction is frequently ignored. Targeting these issues, this paper proposes a novel Temporal Pyramid Transformer (TPT) model with multimodal interaction for VideoQA. The TPT model comprises two modules, namely Question-specific Transformer (QT) and Visual Inference (VI). Given the temporal pyramid constructed from a video, QT builds the question semantics from the coarse-to-fine multimodal co-occurrence between each word and the visual content. Under the guidance of such question-specific semantics, VI infers the visual clues from the local-to-global multi-level interactions between the question and the video. Within each module, we introduce a multimodal attention mechanism to aid the extraction of question-video interactions, with residual connections adopted for the information passing across different levels. Through extensive experiments on three VideoQA datasets, we demonstrate better performances of the proposed method in comparison with the state-of-the-arts.

cs.CV↗

A Joint and Domain-Adaptive Approach to Spoken Language Understanding

Spoken Language Understanding (SLU) is composed of two subtasks: intent detection (ID) and slot filling (SF). There are two lines of research on SLU. One jointly tackles these two subtasks to improve their prediction accuracy, and the other focuses on the domain-adaptation ability of one of the subtasks. In this paper, we attempt to bridge these two lines of research and propose a joint and domain adaptive approach to SLU. We formulate SLU as a constrained generation task and utilize a dynamic vocabulary based on domain-specific ontology. We conduct experiments on the ASMixed and MTOD datasets and achieve competitive performance with previous state-of-the-art joint models. Besides, results show that our joint model can be effectively adapted to a new domain.

cs.CL↗

Deterministic generation of multidimensional photonic cluster states using time-delay feedback

Cluster states are useful in many quantum information processing applications. In particular, universal measurement-based quantum computation (MBQC) utilizes 2D cluster states, and topologically fault-tolerant MBQC requires cluster states with three or higher dimensions. This work proposes a protocol to deterministically generate multidimensional photonic cluster states using a single atom-cavity system and time-delay feedback. The dimensionality of the cluster state increases linearly with the number of time-delay feedback. We firstly give a diagrammatic derivation of the tensor network states, which is valuable in simulating matrix product states and projected entangled pair states generated from sequential photons. Our method also provides a simple way to bridge and analyze the experimental imperfections and the logical errors of the generated states. In this method, we analyze the generated cluster states under realistic experimental conditions and address both one-qubit and two-qubit errors. Through numerical simulation, we observe an optimal atom-cavity cooperativity for the fidelity of the generated states, which is surprising given the prevailing assumption that higher cooperativity systems are inherently better for photonic applications.

quant-ph↗

Transformer-F: A Transformer network with effective methods for learning universal sentence representation

The Transformer model is widely used in natural language processing for sentence representation. However, the previous Transformer-based models focus on function words that have limited meaning in most cases and could merely extract high-level semantic abstraction features. In this paper, two approaches are introduced to improve the performance of Transformers. We calculated the attention score by multiplying the part-of-speech weight vector with the correlation coefficient, which helps extract the words with more practical meaning. The weight vector is obtained by the input text sequence based on the importance of the part-of-speech. Furthermore, we fuse the features of each layer to make the sentence representation results more comprehensive and accurate. In experiments, we demonstrate the effectiveness of our model Transformer-F on three standard text classification datasets. Experimental results show that our proposed model significantly boosts the performance of text classification as compared to the baseline model. Specifically, we obtain a 5.28% relative improvement over the vanilla Transformer on the simple tasks.

cs.CL↗