Searcharxiv⌕ Search

arXiv subjects

Yu Shi

Publications and source records attributed to Yu Shi.

At least 163 records · Page 9Linked to original sources

Exploring the Collective Phenomenon at the Electron-Ion Collider

Based on rare fluctuations in strong interactions, we argue that there is a strong physical resemblance between the high multiplicity events in photo-nuclear collisions and those in $pA$ collisions, in which interesting long range collective phenomena are discovered. This indicates that the collectivity can also be studied in certain kinematic region of the upcoming Electron-Ion Collider (EIC) where the incoming virtual photon has a sufficiently long lifetime. Using a model in the Color Glass Condensate formalism, we first show that the initial state interactions can explain the recent ATLAS azimuthal correlation results measured in the photo-nuclear collisions, and then we provide quantitative predictions for the long range correlations in $eA$ collisions in the EIC regime. With the unprecedented precision and the ability to change the size of the collisional system, the high luminosity EIC will open a new window to explore the physical mechanism responsible for the collective phenomenon.

hep-ph↗

Research on Portfolio Liquidation Strategy under Discrete Times

This paper presents an optimal strategy for portfolio liquidation under discrete time conditions. We assume that N risky assets held will be liquidated according to the same time interval and order quantity, and the basic price processes of assets are generated by an N-dimensional independent standard Brownian motion. The permanent impact generated by an asset in the portfolio during the liquidation will affect all assets, and the temporary impact generated by one asset will only affect itself. On this basis, we establish a liquidation cost model based on the VaR measurement and obtain an optimal liquidation time under discrete-time conditions. The optimal solution shows that the liquidation time is only related to the temporary impact rather than the permanent impact. In the simulation analysis, we give the relationship between volatility parameters, temporary price impact and the optimal liquidation strategy.

q-fin.TR↗

Recognizing Micro-Expression in Video Clip with Adaptive Key-Frame Mining

As a spontaneous expression of emotion on face, micro-expression reveals the underlying emotion that cannot be controlled by human. In micro-expression, facial movement is transient and sparsely localized through time. However, the existing representation based on various deep learning techniques learned from a full video clip is usually redundant. In addition, methods utilizing the single apex frame of each video clip require expert annotations and sacrifice the temporal dynamics. To simultaneously localize and recognize such fleeting facial movements, we propose a novel end-to-end deep learning architecture, referred to as adaptive key-frame mining network (AKMNet). Operating on the video clip of micro-expression, AKMNet is able to learn discriminative spatio-temporal representation by combining spatial features of self-learned local key frames and their global-temporal dynamics. Theoretical analysis and empirical evaluation show that the proposed approach improved recognition accuracy in comparison with state-of-the-art methods on multiple benchmark datasets.

cs.CV↗

Generating Human Readable Transcript for Automatic Speech Recognition with Pre-trained Language Model

Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to disfluency, filter words, and other errata common in spoken communication. Many downstream tasks and human readers rely on the output of the ASR system; therefore, errors introduced by the speaker and ASR system alike will be propagated to the next task in the pipeline. In this work, we propose an ASR post-processing model that aims to transform the incorrect and noisy ASR output into a readable text for humans and downstream tasks. We leverage the Metadata Extraction (MDE) corpus to construct a task-specific dataset for our study. Since the dataset is small, we propose a novel data augmentation method and use a two-stage training strategy to fine-tune the RoBERTa pre-trained model. On the constructed test set, our model outperforms a production two-step pipeline-based post-processing method by a large margin of 13.26 on readability-aware WER (RA-WER) and 17.53 on BLEU metrics. Human evaluation also demonstrates that our method can generate more human-readable transcripts than the baseline method.

cs.CL↗

Improving Zero-shot Neural Machine Translation on Language-specific Encoders-Decoders

Recently, universal neural machine translation (NMT) with shared encoder-decoder gained good performance on zero-shot translation. Unlike universal NMT, jointly trained language-specific encoders-decoders aim to achieve universal representation across non-shared modules, each of which is for a language or language family. The non-shared architecture has the advantage of mitigating internal language competition, especially when the shared vocabulary and model parameters are restricted in their size. However, the performance of using multiple encoders and decoders on zero-shot translation still lags behind universal NMT. In this work, we study zero-shot translation using language-specific encoders-decoders. We propose to generalize the non-shared architecture and universal NMT by differentiating the Transformer layers between language-specific and interlingua. By selectively sharing parameters and applying cross-attentions, we explore maximizing the representation universality and realizing the best alignment of language-agnostic information. We also introduce a denoising auto-encoding (DAE) objective to jointly train the model with the translation task in a multi-task manner. Experiments on two public multilingual parallel datasets show that our proposed model achieves a competitive or better results than universal NMT and strong pivot baseline. Moreover, we experiment incrementally adding new language to the trained model by only updating the new model parameters. With this little effort, the zero-shot translation between this newly added language and existing languages achieves a comparable result with the model trained jointly from scratch on all languages.

cs.CL↗

Speech-language Pre-training for End-to-end Spoken Language Understanding

End-to-end (E2E) spoken language understanding (SLU) can infer semantics directly from speech signal without cascading an automatic speech recognizer (ASR) with a natural language understanding (NLU) module. However, paired utterance recordings and corresponding semantics may not always be available or sufficient to train an E2E SLU model in a real production environment. In this paper, we propose to unify a well-optimized E2E ASR encoder (speech) and a pre-trained language model encoder (language) into a transformer decoder. The unified speech-language pre-trained model (SLP) is continually enhanced on limited labeled data from a target domain by using a conditional masked language model (MLM) objective, and thus can effectively generate a sequence of intent, slot type, and slot value for given input speech in the inference. The experimental results on two public corpora show that our approach to E2E SLU is superior to the conventional cascaded method. It also outperforms the present state-of-the-art approaches to E2E SLU with much less paired data.

cs.CL↗

Underwater Acoustic Multiplexing Communication by Pentamode Metasurface

As the dominant information carrier in water, acoustic wave is widely used for underwater detection, communication and imaging. Even though underwater acoustic communication has been greatly improved in the past decades, it still suffers from the slow transmission speed and low information capacity. The recently developed acoustic orbital angular momentum (OAM) multiplexing communication promises a high efficiency, large capacity and fast transmission speed for acoustic communication. However, the current works on OAM multiplexing communication mainly appears in airborne acoustics. The application of acoustic OAM for underwater communication remains to be further explored and studied. In this paper, an impedance matching pentamode demultiplexing metasurface is designed to realize multiplexing and demultiplexing in underwater acoustic communication. The impedance matching of the metasurface ensures high transmission of the transmitted information. The information encoded into two different OAM beams as two independent channels is numerically demonstrated by realizing real-time picture transfer. The simulation shows the effectiveness of the system for underwater acoustic multiplexing communication. This work paves the way for experimental demonstration and practical application of OAM multiplexing for underwater acoustic communication

physics.app-ph↗

Listen, Look and Deliberate: Visual context-aware speech recognition using pre-trained text-video representations

In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore novel VC-ASR approaches to leverage video and text representations extracted by a self-supervised pre-trained text-video embedding model. Firstly, we propose a multi-stream attention architecture to leverage signals from both audio and video modalities. This architecture consists of separate encoders for the two modalities and a single decoder that attends over them. We show that this architecture is better than fusing modalities at the signal level. Additionally, we also explore leveraging the visual information in a second pass model, which has also been referred to as a `deliberation model'. The deliberation model accepts audio representations and text hypotheses from the first pass ASR and combines them with a visual stream for an improved visual context-aware recognition. The proposed deliberation scheme can work on top of any well trained ASR and also enabled us to leverage the pre-trained text model to ground the hypotheses with the visual features. Our experiments on HOW2 dataset show that multi-stream and deliberation architectures are very effective at the VC-ASR task. We evaluate the proposed models for two scenarios; clean audio stream and distorted audio in which we mask out some specific words in the audio. The deliberation model outperforms the multi-stream model and achieves a relative WER improvement of 6% and 8.7% for the clean and masked data, respectively, compared to an audio-only model. The deliberation model also improves recovering the masked words by 59% relative.

eess.AS↗

The influence of stochastic forcing on strong solutions to the Incompressible Slice Model in 2D bounded domain

The Cotter-Holm Slice Model (CHSM) was introduced to study the behavior of whether and specifically the formulation of atmospheric fronts, whose prediction is fundamental in meteorology. Considered herein is the influence of stochastic forcing on the Incompressible Slice Model (ISM) in a smooth 2D bounded domain, which can be derived by adapting the Lagrangian function in Hamilton's principle for CHSM to the Euler-Boussinesq Eady incompressible case. First, we establish the existence and uniqueness of local pathwise solution (probability strong solution) to the ISM perturbed by nonlinear multiplicative stochastic forcing in Banach spaces $W^{k,p}(D)$ with $k>1+1/p$ and $p\geq 2$. The solution is obtained by introducing suitable cut-off operators applied to the $W^{1,\infty}$-norm of the velocity and temperature fields, using the stochastic compactness method and the Yamada-Watanabe type argument based on the Gyöngy-Krylov characterization of convergence in probability. Then, when the ISM is perturbed by linear multiplicative stochastic forcing and the potential temperature does not vary linearly on the $y$-direction, we prove that the associated Cauchy problem admits a unique global-in-time pathwise solution with high probability, provided that the initial data is sufficiently small or the diffusion parameter is large enough. The results partially answer the problems left open in Alonso-Or{á}n et al. (Physica D 392:99--118, 2019, pp. 117).

math.AP↗

Inverse Designed THz Spectral Splitters

This letter reports proof-of-principle demonstration of 3D printable, low-cost, and compact THz spectral splitters based on diffractive optical elements (DOEs) designed to disperse the incident collimated broadband THz radiation (0.5 THz - 0.7 THz) at a pre-specified distance. Via inverse design, we show that it is possible to design such a diffractive optic, which can split the broadband incident spectrum in any desired fashion, as is evidenced from both FDTD simulations and measured intensity profiles using a 500-750 GHz VNA. Due to its straightforward construction without the usage of any movable parts, our approach, in principle, can have various applications such as in portable, low-cost spectroscopy as well as in wireless THz communication systems as a THz demultiplexer.

physics.optics↗

Mixed-Lingual Pre-training for Cross-lingual Summarization

Cross-lingual Summarization (CLS) aims at producing a summary in the target language for an article in the source language. Traditional solutions employ a two-step approach, i.e. translate then summarize or summarize then translate. Recently, end-to-end models have achieved better results, but these approaches are mostly limited by their dependence on large-scale labeled data. We propose a solution based on mixed-lingual pre-training that leverages both cross-lingual tasks such as translation and monolingual tasks like masked language models. Thus, our model can leverage the massive monolingual data to enhance its modeling of language. Moreover, the architecture has no task-specific components, which saves memory and increases optimization efficiency. We show in experiments that this pre-training scheme can effectively boost the performance of cross-lingual summarization. In Neural Cross-Lingual Summarization (NCLS) dataset, our model achieves an improvement of 2.82 (English to Chinese) and 1.15 (Chinese to English) ROUGE-1 scores over state-of-the-art results.

cs.CL↗

MaP: A Matrix-based Prediction Approach to Improve Span Extraction in Machine Reading Comprehension

Span extraction is an essential problem in machine reading comprehension. Most of the existing algorithms predict the start and end positions of an answer span in the given corresponding context by generating two probability vectors. In this paper, we propose a novel approach that extends the probability vector to a probability matrix. Such a matrix can cover more start-end position pairs. Precisely, to each possible start index, the method always generates an end probability vector. Besides, we propose a sampling-based training strategy to address the computational cost and memory issue in the matrix training phase. We evaluate our method on SQuAD 1.1 and three other question answering benchmarks. Leveraging the most competitive models BERT and BiDAF as the backbone, our proposed approach can get consistent improvements in all datasets, demonstrating the effectiveness of the proposed method.

cs.CL↗

Standard Model of Particle Physics Violating Crypto-Nonlocal Realism

It has been well established that quantum mechanics (QM) violates Bell inequalities (BI), which are consequences of local realism (LR). Remarkably QM also violates Leggett inequalities (LI), which are consequences of a class of nonlocal realism called crypto-nonlocal realism (CNR). Both LR and CNR assume that measurement outcomes are determined by preexisting objective properties, as well as hidden variables (HV) not considered in QM. We extend CNR and LI to include the case that the measurement settings are not externally fixed, but determined by hidden variables (HV). We derive a new version of LI, which is then shown to be violated by entangled $B_d$ mesons, if charge-conjugation-parity (CP) symmetry is indirectly violated, as indeed established. The experimental result is quantitatively estimated by using the indirect CP violation parameter, and the maximum of a suitably defined relative violation is about $2.7\%$. Our work implies that standard model (SM) of particle physics violates CNR. Our LI can also be tested in other systems such as photon polarizations.

hep-ph↗

A Novel Method for Inference of Acyclic Chemical Compounds with Bounded Branch-height Based on Artificial Neural Networks and Integer Programming

Analysis of chemical graphs is a major research topic in computational molecular biology due to its potential applications to drug design. One approach is inverse quantitative structure activity/property relationship (inverse QSAR/QSPR) analysis, which is to infer chemical structures from given chemical activities/properties. Recently, a framework has been proposed for inverse QSAR/QSPR using artificial neural networks (ANN) and mixed integer linear programming (MILP). This method consists of a prediction phase and an inverse prediction phase. In the first phase, a feature vector $f(G)$ of a chemical graph $G$ is introduced and a prediction function $ψ$ on a chemical property $π$ is constructed with an ANN. In the second phase, given a target value $y^*$ of property $π$, a feature vector $x^*$ is inferred by solving an MILP formulated from the trained ANN so that $ψ(x^*)$ is close to $y^*$ and then a set of chemical structures $G^*$ such that $f(G^*)= x^*$ is enumerated by a graph search algorithm. The framework has been applied to the case of chemical compounds with cycle index up to 2. The computational results conducted on instances with $n$ non-hydrogen atoms show that a feature vector $x^*$ can be inferred for up to around $n=40$ whereas graphs $G^*$ can be enumerated for up to $n=15$. When applied to the case of chemical acyclic graphs, the maximum computable diameter of $G^*$ was around up to around 8. We introduce a new characterization of graph structure, "branch-height," based on which an MILP formulation and a graph search algorithm are designed for chemical acyclic graphs. The results of computational experiments using properties such as octanol/water partition coefficient, boiling point and heat of combustion suggest that the proposed method can infer chemical acyclic graphs $G^*$ with $n=50$ and diameter 30.

cs.DS↗

ServiceNet: A P2P Service Network

Given a large number of online services on the Internet, from time to time, people are still struggling to find out the services that they need. On the other hand, when there are considerable research and development on service discovery and service recommendation, most of the related work are centralized and thus suffers inherent shortages of the centralized systems, e.g., adv-driven, lack at trust, transparence and fairness. In this paper, we propose a ServiceNet - a peer-to-peer (P2P) service network for service discovery and service recommendation. ServiceNet is inspired by blockchain technology and aims at providing an open, transparent and self-growth, and self-management service ecosystem. The paper will present the basic idea, an architecture design of the prototype, and an initial implementation and performance evaluation the prototype design.

cs.DC↗

Circuit-based digital adiabatic quantum simulation and pseudoquantum simulation as new approaches to lattice gauge theory

Gauge theory is the framework of the Standard Model of particle physics and is also important in condensed matter physics. As its major non-perturbative approach, lattice gauge theory is traditionally implemented using Monte Carlo simulation, consequently it usually suffers such problems as the Fermion sign problem and the lack of real-time dynamics. Hopefully they can be avoided by using quantum simulation, which simulates quantum systems by using controllable true quantum processes. The field of quantum simulation is under rapid development. Here we present a circuit-based digital scheme of quantum simulation of quantum $\mathbb{Z}_2$ lattice gauge theory in $2+1$ and $3+1$ dimensions, using quantum adiabatic algorithms implemented in terms of universal quantum gates. Our algorithm generalizes the Trotter and symmetric decompositions to the case that the Hamiltonian varies at each step in the decomposition. Furthermore, we carry through a complete demonstration of this scheme in classical GPU simulator, and obtain key features of quantum $\mathbb{Z}_2$ lattice gauge theory, including quantum phase transitions, topological properties, gauge invariance and duality. Hereby dubbed pseudoquantum simulation, classical demonstration of quantum simulation in state-of-art fast computers not only facilitates the development of schemes and algorithms of real quantum simulation, but also represents a new approach of practical computation.

quant-ph↗

Trotter errors in digital adiabatic quantum simulation of quantum $\mathbb{Z}_2$ lattice gauge theory

Trotter decomposition is the basis of the digital quantum simulation. Asymmetric and symmetric decompositions are used in our GPU demonstration of the digital adiabatic quantum simulations of $2+1$ dimensional quantum $\mathbb{Z}_2$ lattice gauge theory. The actual errors in Trotter decompositions are investigated as functions of the coupling parameter and the number of Trotter substeps in each step of the variation of coupling parameter. The relative error of energy is shown to be closely related to the Trotter error usually defined defined in terms of the evolution operators. They are much smaller than the order-of-magnitude estimation. The error in the symmetric decomposition is much smaller than that in the asymmetric decomposition. The features of the Trotter errors obtained here are useful in the experimental implementation of digital quantum simulation and its numerical demonstration.

quant-ph↗

DeepPrognosis: Preoperative Prediction of Pancreatic Cancer Survival and Surgical Margin via Contrast-Enhanced CT Imaging

Pancreatic ductal adenocarcinoma (PDAC) is one of the most lethal cancers and carries a dismal prognosis. Surgery remains the best chance of a potential cure for patients who are eligible for initial resection of PDAC. However, outcomes vary significantly even among the resected patients of the same stage and received similar treatments. Accurate preoperative prognosis of resectable PDACs for personalized treatment is thus highly desired. Nevertheless, there are no automated methods yet to fully exploit the contrast-enhanced computed tomography (CE-CT) imaging for PDAC. Tumor attenuation changes across different CT phases can reflect the tumor internal stromal fractions and vascularization of individual tumors that may impact the clinical outcomes. In this work, we propose a novel deep neural network for the survival prediction of resectable PDAC patients, named as 3D Contrast-Enhanced Convolutional Long Short-Term Memory network(CE-ConvLSTM), which can derive the tumor attenuation signatures or patterns from CE-CT imaging studies. We present a multi-task CNN to accomplish both tasks of outcome and margin prediction where the network benefits from learning the tumor resection margin related features to improve survival prediction. The proposed framework can improve the prediction performances compared with existing state-of-the-art survival analysis approaches. The tumor signature built from our model has evidently added values to be combined with the existing clinical staging system.

eess.IV↗